# Qwen 4: Release Date, Leaks, and the Architecture Alibaba Already Shipped

> Qwen 4 has no date as of Sept 4, 2026, but its architecture is public: Qwen3.8-Flash-Next is the official preview. The confirmed record, the Apsara rumor, the leaks.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-09-04
- Canonical (HTML): https://cellcog.ai/blog/qwen-4-release-date/
- Section: Guides / Choosing a platform
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- Qwen 4 is not released, dated, sized or priced as of September 4, 2026. What Alibaba's Qwen team has done is unusual: on August 26 it open-sourced Qwen3.8-Flash-Next as, in its own words, 'an early preview of the architecture used in Qwen4', released 'before the full Qwen4 model family is built on top of them.'
- The architecture is public and already running. Qwen3.8-Flash-Next is a multimodal mixture-of-experts model with 125B parameters, 6B active per token, a 51B n-gram embedding table and a 4B multi-token-prediction head, 262K native context extensible to 1M, and a claimed training cost around one ninth of Qwen3.7-Plus.
- The inference stack already names it Qwen 4: llama.cpp merged 'model: add Qwen3.8-Flash-Next (qwen4exp)' on August 27; vLLM has open pull requests titled 'Add qwen4 fuse op' and '[Qwen4] Support UVA PLE-offload'; Hugging Face Transformers merged 'Simplify and fix qwen4 tests' on August 26. Code paths exist for a model that has no name on a model card yet.
- The release rumor points at the Apsara Conference, Alibaba Cloud's annual event, scheduled for September 22 to 24 in Hangzhou. Qwen2.5 (2024) and Qwen3-Max (2025) both launched at Apsara; Qwen3 and Qwen3.8-Max did not. No agenda names Qwen 4, and Alibaba's August 20 earnings materials did not mention it.
- One phantom to retire: a 'Qwen 4 Coder 32B-A3B' entry with a June release date and an 82% SWE-bench Verified score circulated in aggregator databases this summer. No such model exists on any Qwen repository; the entry was traced to a data incident and removed.
- This is a living tracker. When Alibaba names a date, a Qwen 4 model card appears, or a checkpoint surfaces, the confirmed facts land here the same day, next to a scorecard of how these claims held up.

## At a glance

- **Is Qwen 4 released?** No. As of September 4, 2026 there is no Qwen 4 model card, weights, API identifier or price. The Qwen team has released the architecture it says Qwen 4 will use, as Qwen3.8-Flash-Next, on August 26, 2026.
- **When is Qwen 4 coming out?** Alibaba has not said. The rumor since July points at the Apsara Conference, September 22 to 24, 2026, because Qwen2.5 and Qwen3-Max both launched there. It is a plausible window with no confirmation; the conference agenda does not name Qwen 4.
- **What has the Qwen team officially said about Qwen 4?** That Qwen3.8-Flash-Next is 'an early preview of the architecture used in Qwen4', released early 'so that the community can examine them before the full Qwen4 model family is built on top of them.' Nothing about timing, sizes or the lineup.
- **Will Qwen 4 be open weights?** Unknown. The architecture preview is open-weight, and Qwen3.8 was the first generation to open a Max-class model (Qwen3.8-2.4T-A95B, August 12). That is a pattern, not a promise; Qwen has said nothing about the Qwen 4 flagship's license.

Qwen 4 is the rare unreleased model whose architecture you can download. On August 26, 2026, Alibaba's Qwen team open-sourced Qwen3.8-Flash-Next and wrote, in the repository's first paragraph, that the model "also serves as an early preview of the architecture used in Qwen4." It added why: "We are again releasing the architectural changes early, so that the community can examine them before the full Qwen4 model family is built on top of them." Nine days later the inference stack has merged code with `qwen4` in its identifiers, and the model those identifiers are for still has no date, no size, no license and no model card.

That gap is what this page tracks. As of September 4, 2026, we pulled the Qwen team's [GitHub release](https://github.com/QwenLM/Qwen3.8-Flash-Next), its [Hugging Face model card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next), Alibaba Cloud's [Apsara Conference page](https://www.alibabacloud.com/en/apsara-conference/2026-about) and Alibaba's [fiscal first-quarter results](https://www.sec.gov/Archives/edgar/data/1577552/000110465926099220/tm2623667d1_ex99-1.htm) directly. Between them, Qwen 4 is named as an architecture and a future family, and nothing else. Every date in circulation is someone's inference.

So this page does what our [Gemini 4 tracker](https://cellcog.ai/blog/gemini-4-release-date/) and our [Qwen3.8-Flash-Next tracker](https://cellcog.ai/blog/qwen3-8-flash-next/) do: every claim dated, official statements separated from reporting and both separated from rumor, updated the same day the story moves. The Flash-Next page is the precedent that matters here: every leaked claim we logged before that release held up on August 26. This one starts with less.

## What we actually know

*Table: Qwen 4 claims vs verifiable status, September 4, 2026*

| Claim | Status |
|---|---|
| Qwen 4 exists as a planned model family | Confirmed by the Qwen team, August 26, 2026 |
| Its architecture is public | Confirmed: Qwen3.8-Flash-Next is the official "early preview of the architecture used in Qwen4" |
| Inference runtimes already implement it | Confirmed: llama.cpp `qwen4exp` merged Aug 27; Transformers `qwen4` tests merged Aug 26; vLLM `[Qwen4]` pull requests open |
| Release date | None from Alibaba; the Apsara Conference (Sep 22 to 24) is the rumored venue, unconfirmed |
| Qwen 4 sizes or lineup (Max, Plus, Flash, Coder) | Nothing official; every figure in circulation is extrapolated from the 125B/6B preview |
| Open weights for the Qwen 4 flagship | Unknown; the preview is open, Qwen3.8 opened a Max-class model, no statement about Qwen 4 |
| "Qwen 4 Coder 32B-A3B, 82% SWE-bench Verified, June" | Phantom; no repository exists, the database entry was removed after reconciliation |
| Ox Alpha on OpenRouter was Qwen 4 | No; revealed as Z.ai's GLM-5.3-Flash on August 26 |
| Qwen 4 benchmarks | None; the only published table is Qwen3.8-Flash-Next's, vendor-reported |
| Alibaba's current flagship | Qwen3.8-Max (August 3), open weights as Qwen3.8-2.4T-A95B (August 12), snapshot 0902 on Code Arena (September 2) |

## What the Qwen team has said, word for word

Two sentences carry the whole official record, both from the August 26 release.

From the [GitHub repository](https://github.com/QwenLM/Qwen3.8-Flash-Next): "In this release we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4." And: "We are again releasing the architectural changes early, so that the community can examine them before the full Qwen4 model family is built on top of them."

From the [Hugging Face model card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next): "This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models interact at scale."

The word "again" is doing work in the first quote. The Qwen team has run this play once before: Qwen3-Next, released in September 2025, previewed the hybrid architecture that the Qwen3.5 family then shipped on, on February 17, 2026. About five months from preview to family. That precedent is the most honest basis anyone has for a Qwen 4 window, and it points at late 2026 or early 2027, not at a conference three weeks away. We will come back to that.

What the team has not said: a month, a size, a lineup, a license, or whether "Qwen 4" will carry the 3.8-Flash-Next configuration up to Max scale or redesign it. Alibaba's own August 20 earnings release, filed with the SEC, describes the quarter's model work as "our flagship foundation model Qwen3.8-Max within three months of its prior version, and we opened its model weights with 2.4 trillion parameters." No Qwen 4.

## The architecture you can already run

The preview is a real model with a published configuration, and it is the only concrete thing anyone can say about Qwen 4's design.

*Table: Qwen3.8-Flash-Next, the official Qwen 4 architecture preview, per the model card*

| Item | Value |
|---|---|
| Released | August 26, 2026, weights on Hugging Face and ModelScope |
| Type | Multimodal mixture-of-experts, text and image in |
| Parameters | 125B main model, 6B active per token, plus 51B n-gram embedding and 4B multi-token-prediction head |
| Context | 262,144 tokens native, extensible to 1,000,000 |
| Attention | Gated DeltaNet hybrid with Qwen Sparse Attention |
| Residual stream | Gated Residual, four branches |
| Embeddings | N-gram Embedding, a lookup table outside the per-token compute |
| Optimizer | Muon, alongside AdamW |
| Training cost | About one ninth of Qwen3.7-Plus, per the Qwen team |
| Production twin | Qwen3.8-Flash on QwenCloud, $0.16 in / $0.47 out per million tokens, 1M context by default |

*Table: Where Qwen3.8-Flash-Next's parameters sit, in billions*

| Component | Billions |
|---|---|
| Main model (MoE) | 125 |
| N-gram embedding table | 51 |
| Multi-token-prediction head | 4 |
| Active per token | 6 |

The design bet is legible from the numbers: park a large share of the parameters in a lookup table that costs almost nothing to read, keep the per-token compute at 6B, and stretch the context with sparse attention. Whether Qwen 4's flagship scales that shape to Max size or pairs it with a dense core is the first thing the model card will tell us, and the thing no leak has credibly claimed.

## The rumor chain, dated

*Table: The Qwen 4 rumor chain, dated*

| Date | Event | Source type |
|---|---|---|
| Jun 2026 | "Qwen 4 Coder 32B-A3B" entry with 82% SWE-bench Verified appears in a model database | Phantom; later removed |
| Jul 7 | Qwen 4 rumored for the September Apsara Conference | Commentator post on X |
| Aug 3 | Qwen3.8-Max released, "the most capable model in the Qwen family to date" | Qwen, official |
| Aug 12 | Qwen3.8-2.4T-A95B weights opened, first Max-class open release | Qwen, official |
| Aug 20 | Alibaba Q1 FY2027 results: Qwen3.8-Max named, no Qwen 4 | Alibaba, official |
| Aug 20 | Ox Alpha appears on OpenRouter; some guess Qwen | Community speculation |
| Aug 26 | Qwen3.8-Flash-Next released as the Qwen 4 architecture preview | Qwen, official |
| Aug 26 | Transformers merges `qwen4` tests; llama.cpp opens `qwen4exp`; vLLM opens `qwen4 fuse op` | Code, public |
| Aug 26 | Ox Alpha revealed as GLM-5.3-Flash, not Qwen | Z.ai, official |
| Aug 30 | Alibaba Cloud posts Apsara 2026 dates: September 22 to 24, Hangzhou | Alibaba Cloud, official |
| Sep 2 | Qwen3.8-Max-0902 snapshot added to Code Arena | Arena changelog |
| Sep 22 to 24 | Apsara Conference 2026 | Scheduled; no Qwen 4 on the agenda |

**The code is the hardest evidence.** Runtimes do not implement an architecture from a rumor. llama.cpp's pull request 27742, titled "model: add Qwen3.8-Flash-Next (qwen4exp)", was opened August 26 and merged August 27. Hugging Face Transformers merged "Simplify and fix qwen4 tests" on August 26 and a "[Qwen4 Exp]" mask fix on September 1. vLLM carries open pull requests titled "[Kernel][Qwen] Add qwen4 fuse op" (August 26), "[Qwen4] Support UVA PLE-offload and N-gram parallelism" (August 29), "[Model][Qwen4] Use ragged PLE IDs for CPU offload" (August 30) and "[Qwen4Exp][ROCm] Support FP8 PLE n-gram checkpoints" (September 3). The maintainers named the architecture what Qwen named it. None of that is a Qwen 4 model; all of it is the ground a Qwen 4 model would land on already being paved.

**The phantom Coder is the cleanest debunk of the summer.** A "Qwen 4 Coder 32B-A3B" entry, complete with a June release date, an Apache 2.0 license and an 82% SWE-bench Verified score, circulated through model databases and the comparison pages that scrape them. No such model exists on the Qwen organizations on Hugging Face, ModelScope or GitHub. The score would have beaten every open-weight coding model on record by a wide margin. The database that carried it, [LLMCheck](https://llmcheck.net/blog/qwen-4-coder-review/), says it ingested the entry during a July data incident and removed it in an August reconciliation against official sources. If a chart shows "Qwen 4 Coder" beating anything, that is its provenance.

**Ox Alpha was never Qwen.** The anonymous OpenRouter model that appeared August 20 drew a Qwen guess from some observers. We traced its fingerprints to Z.ai's GLM family at the time, and Z.ai revealed it as GLM-5.3-Flash on August 26. The full account is on our [Ox Alpha page](https://cellcog.ai/blog/what-is-ox-alpha/). No anonymous checkpoint on Arena or OpenRouter has been credibly attributed to Qwen 4 since.

## Why Apsara keeps coming up, and why it might not be it

The September rumor rests on one real pattern. Alibaba Cloud's Apsara Conference is where Qwen2.5 was revealed on September 19, 2024, and where Qwen3-Max launched on September 24, 2025. This year's conference is [September 22 to 24](https://www.alibabacloud.com/en/apsara-conference/2026-about) at the Hangzhou International Expo Center, dates Alibaba Cloud posted on August 30. A commentator post on X on July 7 put Qwen 4 there, and the idea has circulated since.

Two things cut against it. First, Apsara is not where Qwen generations launch; it is where Qwen ships whatever is ready in late September. Qwen3, the last new generation, launched April 29, 2025, seven months before its conference. Qwen3.8-Max launched August 3, seven weeks before this one. Second, the team's own precedent for an architecture preview says the family follows in months, not weeks: Qwen3-Next in September 2025, the Qwen3.5 family on February 17, 2026. Applied to an August 26 preview, that arithmetic lands in early 2027.

*Table: Where Qwen's recent launches actually happened*

| Model | Date | At Apsara? |
|---|---|---|
| Qwen2.5 | September 19, 2024 | Yes, Apsara 2024 |
| Qwen3 | April 29, 2025 | No |
| Qwen3-Next (architecture preview) | September 2025 | Same week as Apsara 2025 |
| Qwen3-Max | September 24, 2025 | Yes, Apsara 2025 |
| Qwen3.5 (family built on Qwen3-Next) | February 17, 2026 | No |
| Qwen3.8-Max | August 3, 2026 | No |
| Qwen3.8-Flash-Next (Qwen 4 architecture preview) | August 26, 2026 | No |

Our read, labeled as our read: Apsara is the most likely place to hear a Qwen 4 date, and a less likely place to download a Qwen 4 model. A Qwen3.8 refresh on the conference stage, with a Qwen 4 timeline in the keynote, fits every pattern above. A full Qwen 4 family on September 22 would be the fastest preview-to-family turn the team has ever done. We would be glad to be wrong; the page flips either way.

## What shipped instead, and what it costs

The reference point for any Qwen 4 claim is the generation Alibaba is actually shipping, and August was the busiest month the Qwen team has had.

*Table: Qwen3.8 releases, August and September 2026*

| Model | Date | What it is |
|---|---|---|
| Qwen3.8-Max | August 3 | Hosted flagship; "most capable model in the Qwen family to date" |
| Qwen3.8-2.4T-A95B | August 12 | Open weights of the flagship base, 2.4T total, 95B active |
| Qwen3.8-27B | August 14 | Open dense model |
| Qwen3.8-Flash-Next | August 26 | Open Qwen 4 architecture preview, 125B/6B |
| Qwen3.8-Flash | August 28 | Production twin on QwenCloud, $0.16 / $0.47 per million tokens, 1M context |
| Qwen3.8-Max-0902 | September 2 | Post-trained flagship snapshot, on Code Arena |

Six releases in thirty days, two of them open at flagship scale. That cadence is the second reason to read the Qwen 4 rumor calmly: a team shipping this fast does not need Qwen 4 in September to have news at Apsara. It also explains the "within three months of its prior version" line in Alibaba's earnings release. For how the current flagship compares with its closest open rival, see [GLM-5.3 vs Qwen3.8-Max](https://cellcog.ai/blog/glm-5-3-vs-qwen3-8-max/).

We do not route to Qwen today. CellCog's tiers run on Claude Fable 5.1, Claude Opus 5 and Gemini 3.8 Flash, and we moved to each on the day it shipped. Qwen matters to us for a different reason, and it is the reason this page exists: the Qwen 4 architecture preview is open, Qwen3.8 opened a Max-class model for the first time, and if the Qwen 4 flagship follows, it becomes the first frontier-class model an AI employee platform could credibly run without a vendor API. That is a decision about the engine underneath every employee we operate, and we would rather have the facts dated before the day comes.

## What we are watching for

We track open-weight frontier models as an operator weighing engines, not as spectators.

- **A date from Alibaba.** The Apsara keynote on September 22 is the first scheduled chance. A Qwen team post, an Alibaba Cloud release or a model-card commit would count the same.
- **A Qwen 4 model card.** The moment a `Qwen/Qwen4-*` repository appears on Hugging Face or ModelScope, or QwenCloud's release notes list a Qwen 4 identifier, this stops being a rumor page.
- **The license line.** Qwen3.8-Flash-Next shipped under the Qwen Community License. Whether the Qwen 4 flagship is open, open with restrictions, or API-only is the single fact that changes what this model means for self-hosting.
- **The shape at scale.** Does the flagship keep the 125B/6B-with-n-gram-table shape and grow it, or pair the sparse-attention stack with a dense Max-class core? The preview card cannot answer that; the first Qwen 4 card will.
- **Independent replication.** Every Flash-Next number is the Qwen team's. Third-party runs on the preview, and later on Qwen 4, are what make a real comparison possible.
- **Our own routing.** If a Qwen 4 model is real, open and better on our axes, we test it on our infrastructure and this page records the result. If it is not, this page records that too.

## The tracker

This page updates when facts change, not when rumors get louder. As of September 4, 2026: the Qwen 4 architecture is public and running in every major inference engine; the Qwen 4 model is not announced, dated, sized or licensed. The most credible window is not the September conference but the months after it, on the team's own preview-to-family precedent. When Alibaba moves, the confirmed facts and a graded scorecard of the claims above land here the same day.

## Update log

This is a living page; when the story moves, the update lands here.

As of September 4, 2026: page opened. No official movement to log.

## FAQ

**Has Alibaba announced a Qwen 4 release date?**

No. The only official Qwen 4 statements are in the Qwen3.8-Flash-Next release materials of August 26, 2026, which describe an architecture preview and say the full Qwen 4 family will be built on it. Alibaba's fiscal first-quarter results on August 20 discussed Qwen3.8-Max ('within three months of its prior version') and the opened 2.4-trillion-parameter weights, and did not mention Qwen 4. The Apsara Conference page lists September 22 to 24, 2026 and a theme, 'Intelligence Goes Beyond', with no model named.

**What is Qwen3.8-Flash-Next and how is it related to Qwen 4?**

An open-weight multimodal mixture-of-experts model released August 26, 2026, which the Qwen team calls an experimental preview of the architecture that will underpin Qwen 4. Its components: a hybrid of Gated DeltaNet layers and Qwen Sparse Attention, a Gated Residual stream, an N-gram Embedding table of 51B parameters that sits outside the per-token compute budget, and the Muon optimizer. It has 125B parameters with 6B active, plus the 51B table and a 4B multi-token-prediction head, a 262,144-token native context extensible to 1,000,000, and a training cost the team puts at about one ninth of Qwen3.7-Plus. It is not Qwen 4; the product name stays in the 3.8 line. Our full record of that release is on its own page.

**Why does open-source code already say qwen4?**

Because the runtimes implemented the preview under the name the Qwen team gave the architecture. llama.cpp's pull request 27742, 'model: add Qwen3.8-Flash-Next (qwen4exp)', was opened August 26 and merged August 27. Hugging Face Transformers merged 'Simplify and fix qwen4 tests' on August 26 and a '[Qwen4 Exp]' fix on September 1. vLLM has open pull requests titled '[Kernel][Qwen] Add qwen4 fuse op' (August 26), '[Qwen4] Support UVA PLE-offload and N-gram parallelism' (August 29) and '[Model][Qwen4] Use ragged PLE IDs for CPU offload' (August 30). The identifiers mark the architecture generation, not a released model.

**Is Qwen 4 Coder 32B-A3B real?**

No. An entry by that name, with a June 2026 release date, an Apache 2.0 license and an 82% SWE-bench Verified score, appeared in at least one model database this summer and was picked up by comparison sites. There is no such repository on the Qwen organizations on Hugging Face, ModelScope or GitHub, the claimed score would have beaten every open-weight coding model on record, and the database that carried it says it ingested the entry during a July data incident and removed it in an August reconciliation against official sources. If you see 'Qwen 4 Coder' benchmarks, that is where they came from.

**Was the anonymous Ox Alpha model on OpenRouter an early Qwen 4?**

No. Ox Alpha appeared on OpenRouter's stealth endpoint on August 20 with a 1M context and multimodal input, and some observers guessed Qwen. Tokenizer and serving fingerprints pointed elsewhere, and Z.ai revealed it on August 26 as GLM-5.3-Flash. As of September 4 no anonymous Arena or OpenRouter checkpoint has been credibly tied to Qwen 4; the Qwen entry added to Code Arena on September 2 is Qwen3.8-Max-0902, a post-trained snapshot of the current flagship.

**Should I wait for Qwen 4 before building on Qwen?**

No, and the preview is the reason. Anything you run on Qwen3.8-Flash-Next or Qwen3.8-Flash already sits on the Qwen 4 architecture, so the family shipping on top of it is an upgrade to the same stack, not a migration. For our part, CellCog does not route to Qwen today; our tiers run on Claude Fable 5.1, Claude Opus 5 and Gemini 3.8 Flash, and we say so plainly. What a Qwen 4 release changes for an AI employee platform is the open-weight question: a Max-class open model on this architecture would be the first credible self-hosted option at the frontier, and we track that here because the choice of engine sits underneath every employee we run. A complete block of real work on CellCog runs about $25, and the cost depends purely on how much work you assign.

## Related

- [Qwen3.8-Flash-Next Is Out: Confirmed Specs, License, and the Leak Scorecard](https://cellcog.ai/blog/qwen3-8-flash-next/index.md)
- [Qwen3.8-Max-0902: Same Price, Much Better at Coding and Office Work, Still Behind Opus 5](https://cellcog.ai/blog/qwen3-8-max-0902/index.md)
- [Gemini 4: Release Date, Leaks, and What Google Has Actually Said](https://cellcog.ai/blog/gemini-4-release-date/index.md)
- [GLM 5.3 vs Qwen3.8-Max: The Open-Weight Frontier, Compared (August 2026)](https://cellcog.ai/blog/glm-5-3-vs-qwen3-8-max/index.md)

---

Markdown alternate of https://cellcog.ai/blog/qwen-4-release-date/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
