On August 20, 2026, a model with no maker appeared on OpenRouter under the ID stealth/ox-alpha. Three days later, OpenCode’s live data page showed roughly 16 trillion tokens processed, 221,000 unique users, and just over 5 million sessions, making it the #2 model there by recent usage. That was the fastest zero-to-everywhere run of any stealth model this year, and it happened without a single press release, because there was no one to issue it.
Six days in, there was. Update, August 26: the reveal happened, and this page has been updated the same day, as promised.
On this page · 8 sectionsOpen
- RESOLVED: On August 26, 2026, Z.ai confirmed Ox Alpha was a new GLM-series model, and the official listing landed the same day: it is GLM-5.3-Flash, a 320B/18B natively multimodal MoE with MIT-licensed open weights.
- Ox Alpha ran as an anonymous stealth model on OpenRouter (stealth/ox-alpha) from August 20 to 26, 2026: 1,048,576-token context, 131K max output, text, image, and video input, free during the window.
- Adoption exploded: by August 23, OpenCode’s live data showed roughly 16 trillion tokens, 221,000 unique users, and over 5 million sessions, making it the #2 model there by recent usage.
- The public evidence called it before the reveal: a 95-of-95 tokenizer probe match with the GLM-5 vocabulary, Z.ai’s exact API error strings, video-token accounting matching GLM-5V-Turbo, and shared code fingerprints. All four classes were confirmed correct.
- The free window ended with the reveal: the stealth slug is out of OpenRouter’s active catalog, and the official z-ai/glm-5.3-flash route is paid ($0.075/M input and $0.25/M output at the launch promo).
- Still unanswered: what happened to prompts sent during the anonymous window. The routes disagreed on retention, and no post-reveal Z.ai statement addresses the stealth-era data.
- August 30 update: Z.ai’s full engineering write-up officially confirms all stealth-window traffic ran on Chinese AI accelerators, details the 45-layer hybrid-attention design, and publishes the official benchmark table (DeepSWE 63.4 official vs 58.4 community).
§ 01Revealed: it is GLM-5.3-Flash
At 09:00 UTC on August 26, Bloomberg published Z.ai’s confirmation that Ox Alpha was “a new iteration of its GLM series.” By 13:59 UTC, OpenRouter’s production catalog carried the official entry, z-ai/glm-5.3-flash, with the same context window, output ceiling, and modality list the stealth listing had shown all along. The launch blog (“GLM-5.3-Flash: Frontier Intelligence, Flash Cost”), MIT-licensed weights on Hugging Face, and a real price list landed the same day.
What the stealth listing hid: it is a 320B-total, 18B-active mixture-of-experts model, the first natively multimodal member of the GLM-5 series, with a hybrid sparse-plus-linear attention architecture and a 30T-token training corpus. We break down the full specs, the pricing, and the migration path in the reveal post; the pricing page tracks the money side. The rest of this page preserves the stealth-window record, with resolutions marked, because how this model was identified before anyone admitted to it is a story worth keeping.
§ 02What was confirmed during the window
The facts below come from OpenRouter’s live model listing and OpenCode’s documentation as they stood during the stealth window:
| Item | Status then | Status now |
|---|---|---|
| Model ID | stealth/ox-alpha on OpenRouter |
z-ai/glm-5.3-flash (stealth slug delisted) |
| Developer | Undisclosed | Z.ai, confirmed August 26 |
| Context window | 1,048,576 tokens | Unchanged |
| Max output | 131,072 tokens | Unchanged |
| Inputs | Text, images, video | Unchanged |
| Agent features | Tool calling, structured output, reasoning effort | Unchanged |
| Price | $0 preview | $0.075/M in, $0.25/M out (promo); $0.15/$0.50 list |
| Weights | Not released | MIT-licensed, on Hugging Face |
The 1M context was real rather than a metadata placeholder: one researcher’s needle-retrieval tests passed at placements around 934K tokens, failing only above roughly 1.005M.
§ 03The identity mystery, and how it was solved
Nobody claimed Ox Alpha during the window, so everything about its origin was inference. The inference turned out to be excellent. Four independent evidence classes pointed at Z.ai’s GLM family, and the reveal confirmed all four:
The tokenizer. An investigation published August 22 and updated August 23 ran 95 discriminating probe strings and reported a 95-of-95 exact match with the GLM-5-generation vocabulary, mean absolute error zero, plus a chat-template grammar that matched Z.ai’s served GLM-5.2 endpoint down to its malformed-input behavior. Confirmed.
The error strings. Malformed requests returned Z.ai’s exact error envelope, code 1214, “Incorrect role information,” and one malformed parameter exposed an internal Java class path from the serving stack. This class of evidence pointed at the operator, not just the model family. Confirmed: Z.ai was serving its own model.
The video encoder. Controlled video inputs produced token budgets matching GLM-5V-Turbo’s known accounting. Encoders are much harder to match by accident than prose style, and the revealed model is indeed natively multimodal GLM. Confirmed in direction.
The code style. An independent seven-model comparison found Ox Alpha and GLM-5.3 uniquely sharing rare fingerprints: an identical helper function, the same ORM pairing, the same habit of rewriting the README with identical structure. Confirmed: same family, and per Z.ai the Flash model starts from a newly trained base within it.
The competing theory, an unreleased Gemini checkpoint fueled by DeepMind researchers’ hints on X, had no comparable technical evidence and was wrong. The durable lesson for the next stealth model, and there will be one: tokenizers and error strings do not know how to lie, and a model naming itself in chat proves nothing in either direction.
§ 04Was it actually good?
The viral number said 80 percent. The fine print said that was 8 of 10 hand-picked coding tasks in one independent test. The completed run told the real story: a full 113-task DeepSWE attempt finished on August 23 with 66 tasks resolved, a 58.4 percent resolve rate, one attempt per task, with 80 percent of tasks passing at least 90 percent of their tests and roughly a tenth of the run lost to tool-call formatting rather than reasoning. A separate LiveCodeBench v6 run scored 28 percent Pass@1 in a direct-generation setup with no agent harness, which mostly confirmed that this model wants tools and a terminal, not a chat box.
Z.ai’s official card now stakes bigger claims: that GLM-5.3-Flash outperforms GLM-5.2 across benchmarks at one-tenth the price, and approaches Claude Opus 4.8 on coding and agentic benchmarks. Those are the vendor’s numbers and carry the vendor’s framing, but the shape matches what the stealth window showed independently: genuinely strong long-horizon agentic work. The most impressive reports clustered exactly there, like the user who migrated two large projects across a 100K-to-200K-token working set that had forced constant context compaction in other tools. Users who put it in a real agent harness consistently rated it higher than users who just chatted with it.
§ 05The data question, then and now
During the window, Ox Alpha was free, anonymous, and temporary, and we flagged the retention contradiction as the most informative fine print on the model: OpenCode said the provider followed zero retention, while OpenRouter’s own stealth listing said prompts and completions were retained by the provider. Both could not be the whole story.
The reveal resolves half of it. The counterparty now has a name, published terms, and a reputation to protect, which is a categorical improvement over an anonymous box. What it does not resolve: which retention story was true during the window, and what happens to the prompts sent before August 26. We have found no Z.ai statement addressing the stealth-era data specifically. If you sent it credentials, customer data, or sensitive proprietary code during the free window, that data sits under whatever the stealth arrangement actually was, and nobody has said what that arrangement is.
§ 06What Ox Alpha was not
It was a model, not an agent, and the distinction is where all the practical decisions lived, both during the window and after it. Ox Alpha accepted inputs and emitted text or tool calls. Everything that made that useful for real work, the tools, filesystem and browser access, permissions, retries, verification, memory that survives the session, came from the harness around it. That is why its heaviest usage ran inside coding harnesses rather than chat interfaces, and it is the same reason we argued in our harness ranking that the harness layer, not the model layer, is where working AI actually gets decided.
That layer is also where we live, so discount accordingly: CellCog is the employee layer above models like this one. Models rotate under our platform as they leapfrog each other, and the point of a standing AI employee is that the rotation does not matter to you. The employee keeps its role, its context, and its continuity; the model underneath is an implementation detail that gets better every few months, sometimes via a mystery box on OpenRouter that turns out to be a 320B-parameter open-weight model wearing a codename.
§ 07How it ended
Of the three outcomes we tracked, the first one hit, and on schedule: a lab claimed it, pricing appeared, and the free-window users became launch-day customers, which was very likely the strategy from the start. The window ran August 20 to 26, one day inside the community’s August 27 estimate. The full reveal record, specs, and migration path live in the GLM-5.3-Flash post, updated as the story develops.
§ 08Update, August 30: the official engineering write-up
Four days after the reveal, Z.ai published its complete technical article on GLM-5.3-Flash: “More Intelligence with Less Compute,” dated August 31 China time, which is August 30 in the US. Three parts of it bear directly on the record this page keeps.
The hardware is now official. The write-up states it in one sentence: “we evaluated GLM-5.3-Flash anonymously as Ox Alpha under real-world traffic, with all traffic served on Chinese AI accelerators.” The stealth window was a hardware demonstration as much as a model one. Six days of free million-token-context inference for 221,000 users ran at production scale on domestic Chinese accelerators, and Z.ai wants that fact in the launch record. Given the export-control backdrop, that sentence may age into the most consequential line in the post.
The efficiency story has its mechanism. The architecture section explains how a 320B model serves this cheaply: 45 layers against the GLM-4.5 series’ 92, active parameters cut from 32B to 18B, linear attention handling local dependencies while a lightweight sparse-attention indexer retrieves from the global context, and a compression scheme called IndexPool that pools four cached key vectors into one at long context. Z.ai’s comparison claims roughly 3.0x less attention compute and a 4.4x smaller KV cache than GLM-5.3, with one honest caveat included: the KV-cache footprint still runs slightly larger than some comparable Flash-class models.
The official benchmarks arrived, and the community number holds up. The write-up publishes an eight-row table against GLM-5.2, including rows the reveal-day card did not show: NL2Repo at 56.3 and Toolathlon Verified at 78.4. The official DeepSWE v1.1 figure is 63.4, against the 58.4 that the completed community run measured during the stealth window. Both are real numbers from different harnesses: Z.ai’s footnoted setup is more generous than the community’s one-attempt run. The full table and analysis live in the reveal post.
What the write-up still does not address, for the record: the stealth-era data question. Nothing in the article says what happened to prompts sent during the anonymous window. The open question above stays open.
Q1What was Ox Alpha revealed to be?
GLM-5.3-Flash from Z.ai: the first natively multimodal model in the GLM-5 series, a 320B-total, 18B-active mixture-of-experts with hybrid sparse-plus-linear attention, trained on a 30T-token corpus, released with MIT-licensed weights on Hugging Face on August 26, 2026. Our full reveal breakdown covers the specs and pricing.
Q2Was the GLM theory right all along?
Yes, on every count. The 95-of-95 tokenizer probe match, Z.ai’s exact error envelopes (code 1214), the GLM-5V-style video-token accounting, and the shared code fingerprints all pointed at the GLM family, and the reveal confirmed each class. The competing Gemini theory, based on researcher hints, was wrong.
Q3What were Ox Alpha's specs?
Per OpenRouter’s listing, identical before and after the reveal: 1,048,576-token context, 131,072-token max output, text, image, and video input, text output, mandatory reasoning with effort settings, tool calling, and structured output. The reveal added what the stealth listing hid: 320B total parameters, 18B active, and open MIT weights.
Q4Can I still use the stealth/ox-alpha endpoint?
No. The slug is out of OpenRouter’s active catalog, no alias is documented, and the free preview pricing is gone. Point code at z-ai/glm-5.3-flash instead, or self-host the MIT weights with SGLang, vLLM, TokenSpeed, or KTransformers.
Q5Can I trust it with my code now?
The counterparty question is resolved: it is Z.ai, a named company with published terms. What remains unresolved is the stealth window itself. The routes disagreed on retention while the model was anonymous, and we have found no Z.ai statement about what happened to prompts sent during that period. Data you sent before August 26 sits under whatever the stealth arrangement actually was.
Q6What did the stealth window turn out to be?
Launch marketing, almost certainly by design. The Hugging Face repository was created a day before the reveal, the free window built a six-figure user base at breathtaking speed, and the users burning the free tokens hardest became the launch-day audience. Expect the pattern again: it worked.
