Ox Alpha’s full price list, for the six days it existed under that name, was one number: $0. On August 26, 2026, the number changed, because the model got a name: Z.ai revealed it as GLM-5.3-Flash, shipped MIT-licensed weights, and published a real price list. This page updated the same day, as promised. The full reveal record, specs and all, is in the GLM-5.3-Flash post; the stealth-window story is preserved in the explainer. This page keeps doing its job: tracking the money.
- RESOLVED: The free window ended on August 26, 2026, the day Z.ai revealed Ox Alpha as GLM-5.3-Flash. The stealth slug is out of OpenRouter’s active catalog and the official route is paid.
- Real pricing, per 1M tokens: Z.ai lists $0.15 input and $0.50 output, and the 50 percent launch promo ($0.075 and $0.25) ended September 9, 2026. OpenRouter’s z-ai/glm-5.3-flash route still charged the promo numbers on September 9, the first gap between the two routes.
- The GLM-family prediction held: post-reveal pricing landed in the aggressive open-weight tier, roughly one-tenth of the GLM-5.3 flagship’s $1.40/$4.40, nowhere near frontier-flagship rates.
- A third option replaced the free tier: MIT-licensed weights on Hugging Face with day-one SGLang, vLLM, TokenSpeed, and KTransformers support. Self-hosting a 320B MoE is a datacenter decision, not a laptop one.
- The retention contradiction between routes was never resolved for the stealth window itself: no Z.ai statement addresses what happened to prompts sent before the reveal.
- This page updated the same day the pricing landed and again on September 9, the day the promo ended, as promised.
§ 01Every route, now
| Route | Input | Output | Cached input | Fine print |
|---|---|---|---|---|
OpenRouter (z-ai/glm-5.3-flash) |
$0.075 | $0.25 | $0.015 | Still the promo numbers at 19:00 UTC Sept 9, now half of Z.ai’s list; stealth slug delisted |
| Z.ai API, launch promo (ended) | $0.075 | $0.25 | $0.015 | 50 percent discount, ended 24:00 Sept 9, 2026 (UTC+8) |
| Z.ai API, list | $0.15 | $0.50 | $0.03 | The current first-party price; cache storage still limited-time free |
| GLM Coding Plan | Subscription | Subscription | n/a | “Fully available” with 3x quota, per Z.ai’s docs |
| Self-hosted | Your hardware | Your hardware | n/a | MIT weights on Hugging Face; 320B MoE, plan for datacenter GPUs |
Two things about this table would have been worth money a week ago. First, on reveal day the launch promo and the OpenRouter price were the same number, so there was no arbitrage between routes; that stopped being true on September 9, when Z.ai moved to list and OpenRouter did not, and the update below has the numbers. Second, the price landed exactly where the GLM-family theory said it would: in that family’s aggressive range, roughly a tenth of the GLM-5.3 flagship’s $1.40 input and $4.40 output, and nowhere near frontier-flagship rates. The identity evidence predicted the invoice.
§ 02Update, September 9: the promo ended, and the routes split
As promised, here is the note. The 50 percent launch discount expired at 24:00 Singapore time on September 9, and Z.ai’s pricing page now shows GLM-5.3-Flash at list only, $0.15 per million input, $0.03 cached input and $0.50 output, with the promo sentence removed. Three hours later OpenRouter had not followed: the endpoint OpenRouter labels Z.AI and the DeepInfra endpoint both still quoted $0.075, $0.25 and $0.015 at 19:00 UTC. So there is arbitrage between routes for the first time since the reveal, and it points at the aggregator. We will update this section when it closes.
| Route | Input | Output | Cached input | Note |
|---|---|---|---|---|
| Z.ai API, list | $0.15 | $0.50 | $0.03 | Promo line removed from the pricing page |
| OpenRouter, Z.AI endpoint | $0.075 | $0.25 | $0.015 | Still the launch-promo numbers |
| OpenRouter, DeepInfra | $0.075 | $0.25 | $0.015 | Same as the Z.AI endpoint |
Update, September 11: Z.ai’s own OpenRouter endpoint moved to list ($0.15 in, $0.50 out) by September 11, 2026; DeepInfra ($0.075 and $0.25) and Relace ($0.09 and $0.30) still undercut it, so the cheapest paid route for the former Ox Alpha is now a third-party host on OpenRouter, not Z.ai. Read 08:28 UTC. | OpenRouter, Wafer | $0.10 | $0.35 | $0.02 | Third-party host | | OpenRouter, GMICloud | $0.1125 | $0.375 | $0.0225 | Third-party host |
Alongside it, Z.ai is running a GLM Coding Plan usage campaign from September 3 to September 20, 2026, 23:00 to 09:00 UTC+8 daily, zero quota consumption in ZCode and doubled quota in other supported agents for paid Coding Plan users, per Z.ai’s notice; and the MIT-licensed weights now have community quantizations on Hugging Face with six-figure download counts, the one price that never expires.
§ 03How the window actually ended
The one real date signal during the stealth run was OpenCode’s August 20 launch note promising one free week, which pointed at roughly August 27. The window beat the estimate by a day: it ended August 26, not with a shutdown but with a reveal, which was always the likely exit. The endpoint did not vanish; it got a name, a license, and a meter. The stealth slug stealth/ox-alpha is out of OpenRouter’s active catalog with no documented alias, so the practical deadline for anyone still pointed at it is immediate: switch the model ID to z-ai/glm-5.3-flash, self-host the MIT weights, or stop calling it.
What the $0 bought the provider is now legible too. OpenCode’s counters showed roughly 16 trillion tokens processed and 221,000 users inside the window, and the Hugging Face repository was created the day before the reveal. The free week was launch marketing with the launch list pre-assembled, executed cleanly. Expect the pattern to be copied.
§ 04What free cost, in the end
The retention contradiction we flagged during the window never got resolved, it just got a counterparty. OpenCode said zero retention; OpenRouter’s stealth listing said the provider retains prompts and completions. Both could not be the whole story, and Z.ai has not, as of the reveal, published anything that says which one was. Prompts sent during the anonymous window sit under terms nobody has stated. That was the real price of free, and for teams that kept sensitive code out of it, the window was exactly what it looked like: a genuinely generous evaluation period for a genuinely strong model.
If you are budgeting around the model now, the practical rule survives the reveal unchanged: benchmark it on your real workloads, and write the list prices into the budget; September 9 arrived, as deadlines do, faster than a migration.
§ 05The cost frame that survived the reveal
Here is the comparison we said would still be true after the window closed, and it is. GLM-5.3-Flash is a model, priced per token, and now priced honestly and aggressively. The things businesses actually budget for, a role that gets owned, work that carries over, someone accountable between conversations, are a layer above the token price, and that layer is where the real cost differences live. Token prices across strong models have been converging for a year; the reveal just added another datapoint at the cheap end of the curve. What has not converged is what sits on top of them.
That is the layer we build, so apply the usual discount: CellCog sells standing AI employees, usage-based, where you pay for the work, not the hire, and the cost depends purely on how much work you assign. The model underneath is our problem, not yours, including on the day a mystery box on OpenRouter turns out to be a 320B-parameter open-weight model with a price list. Our full cost breakdown, worked scenarios included, is in the AI employee cost guide.
Q1What does GLM-5.3-Flash (Ox Alpha) cost on OpenRouter?
As of September 9, 2026, 19:00 UTC: still $0.075 per million input tokens, $0.25 per million output, and $0.015 per million cached input, under the model ID z-ai/glm-5.3-flash. Those were Z.ai’s launch-promo numbers; Z.ai’s own API moved to list on September 9 and OpenRouter had not followed yet. The old stealth/ox-alpha slug is out of the active catalog.
Q2What will it cost after the launch promo?
Z.ai’s list prices: $0.15 per million input, $0.50 per million output, $0.03 per million cached input. The 50 percent discount ended at 24:00 on September 9, 2026, Singapore time, and those are now the only first-party numbers on the pricing page. Budget against list.
Q3Is Ox Alpha still free anywhere?
No route we can verify still serves it free. The official OpenRouter route is paid, and the OpenCode promotion was explicitly a limited-time preview offer that the reveal concluded. The lasting free option is different in kind: the weights themselves are MIT-licensed and self-hostable.
Q4How does the price compare to other models?
At list, GLM-5.3-Flash costs roughly one-tenth of the GLM-5.3 flagship ($1.40 input, $4.40 output), which is exactly the ratio Z.ai markets. That lands it in the aggressive open-weight tier, as the GLM-family theory predicted, rather than anywhere near frontier-flagship rates.
Q5Did using free Ox Alpha cost anything in the end?
Money, no. The open question is data: the access routes stated contradictory retention terms during the anonymous window, and no post-reveal statement from Z.ai addresses what happened to prompts sent before August 26. Sensitive code sent during the window sits under terms nobody has published.
