Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact

Qwen 4: Release Date, Leaks, and the Architecture Alibaba Already Shipped

At a glanceQuick answers
Is Qwen 4 released?
No. As of September 4, 2026 there is no Qwen 4 model card, weights, API identifier or price. The Qwen team has released the architecture it says Qwen 4 will use, as Qwen3.8-Flash-Next, on August 26, 2026.
When is Qwen 4 coming out?
Alibaba has not said. The rumor since July points at the Apsara Conference, September 22 to 24, 2026, because Qwen2.5 and Qwen3-Max both launched there. It is a plausible window with no confirmation; the conference agenda does not name Qwen 4.
What has the Qwen team officially said about Qwen 4?
That Qwen3.8-Flash-Next is ‘an early preview of the architecture used in Qwen4’, released early ‘so that the community can examine them before the full Qwen4 model family is built on top of them.’ Nothing about timing, sizes or the lineup.
Will Qwen 4 be open weights?
Unknown. The architecture preview is open-weight, and Qwen3.8 was the first generation to open a Max-class model (Qwen3.8-2.4T-A95B, August 12). That is a pattern, not a promise; Qwen has said nothing about the Qwen 4 flagship’s license.
Hand-drawn sketch of an unrolled blueprint labeled QWEN4 ARCHITECTURE stamped OPEN AUG 26, a small solid engine labeled 3.8 FLASH-NEXT built from it, a much larger dashed outline of an engine labeled QWEN 4, and a conference badge reading APSARA SEP 22-24 with an amber question mark
Fig 0The blueprint is public and one engine has already been built from it. The big one is still a dashed line, and the badge has a question mark where the date should be.

Qwen 4 is the rare unreleased model whose architecture you can download. On August 26, 2026, Alibaba’s Qwen team open-sourced Qwen3.8-Flash-Next and wrote, in the repository’s first paragraph, that the model “also serves as an early preview of the architecture used in Qwen4.” It added why: “We are again releasing the architectural changes early, so that the community can examine them before the full Qwen4 model family is built on top of them.” Nine days later the inference stack has merged code with qwen4 in its identifiers, and the model those identifiers are for still has no date, no size, no license and no model card.

That gap is what this page tracks. As of September 4, 2026, we pulled the Qwen team’s GitHub release, its Hugging Face model card, Alibaba Cloud’s Apsara Conference page and Alibaba’s fiscal first-quarter results directly. Between them, Qwen 4 is named as an architecture and a future family, and nothing else. Every date in circulation is someone’s inference.

So this page does what our Gemini 4 tracker and our Qwen3.8-Flash-Next tracker do: every claim dated, official statements separated from reporting and both separated from rumor, updated the same day the story moves. The Flash-Next page is the precedent that matters here: every leaked claim we logged before that release held up on August 26. This one starts with less.

On this page · 9 sectionsOpen
  1. What we actually know
  2. What the Qwen team has said, word for word
  3. The architecture you can already run
  4. The rumor chain, dated
  5. Why Apsara keeps coming up, and why it might not be it
  6. What shipped instead, and what it costs
  7. What we are watching for
  8. The tracker
  9. Update log
Key points6 · 15 min full read
  1. Qwen 4 is not released, dated, sized or priced as of September 4, 2026. What Alibaba’s Qwen team has done is unusual: on August 26 it open-sourced Qwen3.8-Flash-Next as, in its own words, ‘an early preview of the architecture used in Qwen4’, released ‘before the full Qwen4 model family is built on top of them.’
  2. The architecture is public and already running. Qwen3.8-Flash-Next is a multimodal mixture-of-experts model with 125B parameters, 6B active per token, a 51B n-gram embedding table and a 4B multi-token-prediction head, 262K native context extensible to 1M, and a claimed training cost around one ninth of Qwen3.7-Plus.
  3. The inference stack already names it Qwen 4: llama.cpp merged ‘model: add Qwen3.8-Flash-Next (qwen4exp)’ on August 27; vLLM has open pull requests titled ‘Add qwen4 fuse op’ and ‘[Qwen4] Support UVA PLE-offload’; Hugging Face Transformers merged ‘Simplify and fix qwen4 tests’ on August 26. Code paths exist for a model that has no name on a model card yet.
  4. The release rumor points at the Apsara Conference, Alibaba Cloud’s annual event, scheduled for September 22 to 24 in Hangzhou. Qwen2.5 (2024) and Qwen3-Max (2025) both launched at Apsara; Qwen3 and Qwen3.8-Max did not. No agenda names Qwen 4, and Alibaba’s August 20 earnings materials did not mention it.
  5. One phantom to retire: a ‘Qwen 4 Coder 32B-A3B’ entry with a June release date and an 82% SWE-bench Verified score circulated in aggregator databases this summer. No such model exists on any Qwen repository; the entry was traced to a data incident and removed.
  6. This is a living tracker. When Alibaba names a date, a Qwen 4 model card appears, or a checkpoint surfaces, the confirmed facts land here the same day, next to a scorecard of how these claims held up.

§ 01What we actually know

Claim Status
Qwen 4 exists as a planned model family Confirmed by the Qwen team, August 26, 2026
Its architecture is public Confirmed: Qwen3.8-Flash-Next is the official “early preview of the architecture used in Qwen4”
Inference runtimes already implement it Confirmed: llama.cpp qwen4exp merged Aug 27; Transformers qwen4 tests merged Aug 26; vLLM [Qwen4] pull requests open
Release date None from Alibaba; the Apsara Conference (Sep 22 to 24) is the rumored venue, unconfirmed
Qwen 4 sizes or lineup (Max, Plus, Flash, Coder) Nothing official; every figure in circulation is extrapolated from the 125B/6B preview
Open weights for the Qwen 4 flagship Unknown; the preview is open, Qwen3.8 opened a Max-class model, no statement about Qwen 4
“Qwen 4 Coder 32B-A3B, 82% SWE-bench Verified, June” Phantom; no repository exists, the database entry was removed after reconciliation
Ox Alpha on OpenRouter was Qwen 4 No; revealed as Z.ai’s GLM-5.3-Flash on August 26
Qwen 4 benchmarks None; the only published table is Qwen3.8-Flash-Next’s, vendor-reported
Alibaba’s current flagship Qwen3.8-Max (August 3), open weights as Qwen3.8-2.4T-A95B (August 12), snapshot 0902 on Code Arena (September 2)
Table 1Qwen 4 claims vs verifiable status, September 4, 2026

§ 02What the Qwen team has said, word for word

Two sentences carry the whole official record, both from the August 26 release.

From the GitHub repository: “In this release we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4.” And: “We are again releasing the architectural changes early, so that the community can examine them before the full Qwen4 model family is built on top of them.”

From the Hugging Face model card: “This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models interact at scale.”

The word “again” is doing work in the first quote. The Qwen team has run this play once before: Qwen3-Next, released in September 2025, previewed the hybrid architecture that the Qwen3.5 family then shipped on, on February 17, 2026. About five months from preview to family. That precedent is the most honest basis anyone has for a Qwen 4 window, and it points at late 2026 or early 2027, not at a conference three weeks away. We will come back to that.

What the team has not said: a month, a size, a lineup, a license, or whether “Qwen 4” will carry the 3.8-Flash-Next configuration up to Max scale or redesign it. Alibaba’s own August 20 earnings release, filed with the SEC, describes the quarter’s model work as “our flagship foundation model Qwen3.8-Max within three months of its prior version, and we opened its model weights with 2.4 trillion parameters.” No Qwen 4.

§ 03The architecture you can already run

The preview is a real model with a published configuration, and it is the only concrete thing anyone can say about Qwen 4’s design.

Item Value
Released August 26, 2026, weights on Hugging Face and ModelScope
Type Multimodal mixture-of-experts, text and image in
Parameters 125B main model, 6B active per token, plus 51B n-gram embedding and 4B multi-token-prediction head
Context 262,144 tokens native, extensible to 1,000,000
Attention Gated DeltaNet hybrid with Qwen Sparse Attention
Residual stream Gated Residual, four branches
Embeddings N-gram Embedding, a lookup table outside the per-token compute
Optimizer Muon, alongside AdamW
Training cost About one ninth of Qwen3.7-Plus, per the Qwen team
Production twin Qwen3.8-Flash on QwenCloud, $0.16 in / $0.47 out per million tokens, 1M context by default
Table 2Qwen3.8-Flash-Next, the official Qwen 4 architecture preview, per the model card
Component Billions
Main model (MoE) 125
N-gram embedding table 51
Multi-token-prediction head 4
Active per token 6
Table 3Where Qwen3.8-Flash-Next’s parameters sit, in billions
Qwen3.8-Flash-Next parameter budget, in billionsBar chart of the preview model's parameter budget from its model card: main MoE 125 billion, n-gram embedding table 51 billion, multi-token-prediction head 4 billion, and 6 billion active per token highlightedMain MoE125N-gram table51MTP head4Active per token6Qwen3.8-Flash-Next parameter budget, in billionsBar chart of the preview model's parameter budget from its model card: main MoE 125 billion, n-gram embedding table 51 billion, multi-token-prediction head 4 billion, and 6 billion active per token highlightedMain MoE125N-gram table51MTP head4Active per token6
Fig 1Qwen3.8-Flash-Next parameter budget, in billions

The design bet is legible from the numbers: park a large share of the parameters in a lookup table that costs almost nothing to read, keep the per-token compute at 6B, and stretch the context with sparse attention. Whether Qwen 4’s flagship scales that shape to Max size or pairs it with a dense core is the first thing the model card will tell us, and the thing no leak has credibly claimed.

§ 04The rumor chain, dated

Date Event Source type
Jun 2026 “Qwen 4 Coder 32B-A3B” entry with 82% SWE-bench Verified appears in a model database Phantom; later removed
Jul 7 Qwen 4 rumored for the September Apsara Conference Commentator post on X
Aug 3 Qwen3.8-Max released, “the most capable model in the Qwen family to date” Qwen, official
Aug 12 Qwen3.8-2.4T-A95B weights opened, first Max-class open release Qwen, official
Aug 20 Alibaba Q1 FY2027 results: Qwen3.8-Max named, no Qwen 4 Alibaba, official
Aug 20 Ox Alpha appears on OpenRouter; some guess Qwen Community speculation
Aug 26 Qwen3.8-Flash-Next released as the Qwen 4 architecture preview Qwen, official
Aug 26 Transformers merges qwen4 tests; llama.cpp opens qwen4exp; vLLM opens qwen4 fuse op Code, public
Aug 26 Ox Alpha revealed as GLM-5.3-Flash, not Qwen Z.ai, official
Aug 30 Alibaba Cloud posts Apsara 2026 dates: September 22 to 24, Hangzhou Alibaba Cloud, official
Sep 2 Qwen3.8-Max-0902 snapshot added to Code Arena Arena changelog
Sep 22 to 24 Apsara Conference 2026 Scheduled; no Qwen 4 on the agenda
Table 4The Qwen 4 rumor chain, dated
From a phantom to a public architecture, with a conference at the endTimeline from the July 7 Apsara rumor through the August 3 Qwen3.8-Max release, the August 12 open weights, the August 20 earnings call with no Qwen 4, the August 26 Flash-Next architecture preview highlighted, the August 30 Apsara date announcement, and the September 22 conference the rumor points atJul 7Rumor: Qwen 4 at ApsaraAug 3Qwen3.8-Max releasedAug 122.4T-A95B weights openedAug 20Earnings: Qwen3.8-Max, no Qwen 4Aug 26Flash-Next: Qwen 4 architecture previewAug 30Apsara 2026 dated: Sep 22 to 24Sep 22Apsara Conference opensFrom a phantom to a public architecture, with a conference at the endTimeline from the July 7 Apsara rumor through the August 3 Qwen3.8-Max release, the August 12 open weights, the August 20 earnings call with no Qwen 4, the August 26 Flash-Next architecture preview highlighted, the August 30 Apsara date announcement, and the September 22 conference the rumor points atJul 7Rumor: Qwen 4 at ApsaraAug 3Qwen3.8-Max releasedAug 122.4T-A95B weights openedAug 20Earnings: Qwen3.8-Max, no Qwen 4Aug 26Flash-Next: Qwen 4 architecture previewAug 30Apsara 2026 dated: Sep 22 to 24Sep 22Apsara Conference opens
Fig 2From a phantom to a public architecture, with a conference at the end

The code is the hardest evidence. Runtimes do not implement an architecture from a rumor. llama.cpp’s pull request 27742, titled “model: add Qwen3.8-Flash-Next (qwen4exp)”, was opened August 26 and merged August 27. Hugging Face Transformers merged “Simplify and fix qwen4 tests” on August 26 and a “[Qwen4 Exp]” mask fix on September 1. vLLM carries open pull requests titled “[Kernel][Qwen] Add qwen4 fuse op” (August 26), “[Qwen4] Support UVA PLE-offload and N-gram parallelism” (August 29), “[Model][Qwen4] Use ragged PLE IDs for CPU offload” (August 30) and “[Qwen4Exp][ROCm] Support FP8 PLE n-gram checkpoints” (September 3). The maintainers named the architecture what Qwen named it. None of that is a Qwen 4 model; all of it is the ground a Qwen 4 model would land on already being paved.

The phantom Coder is the cleanest debunk of the summer. A “Qwen 4 Coder 32B-A3B” entry, complete with a June release date, an Apache 2.0 license and an 82% SWE-bench Verified score, circulated through model databases and the comparison pages that scrape them. No such model exists on the Qwen organizations on Hugging Face, ModelScope or GitHub. The score would have beaten every open-weight coding model on record by a wide margin. The database that carried it, LLMCheck, says it ingested the entry during a July data incident and removed it in an August reconciliation against official sources. If a chart shows “Qwen 4 Coder” beating anything, that is its provenance.

Ox Alpha was never Qwen. The anonymous OpenRouter model that appeared August 20 drew a Qwen guess from some observers. We traced its fingerprints to Z.ai’s GLM family at the time, and Z.ai revealed it as GLM-5.3-Flash on August 26. The full account is on our Ox Alpha page. No anonymous checkpoint on Arena or OpenRouter has been credibly attributed to Qwen 4 since.

§ 05Why Apsara keeps coming up, and why it might not be it

The September rumor rests on one real pattern. Alibaba Cloud’s Apsara Conference is where Qwen2.5 was revealed on September 19, 2024, and where Qwen3-Max launched on September 24, 2025. This year’s conference is September 22 to 24 at the Hangzhou International Expo Center, dates Alibaba Cloud posted on August 30. A commentator post on X on July 7 put Qwen 4 there, and the idea has circulated since.

Two things cut against it. First, Apsara is not where Qwen generations launch; it is where Qwen ships whatever is ready in late September. Qwen3, the last new generation, launched April 29, 2025, seven months before its conference. Qwen3.8-Max launched August 3, seven weeks before this one. Second, the team’s own precedent for an architecture preview says the family follows in months, not weeks: Qwen3-Next in September 2025, the Qwen3.5 family on February 17, 2026. Applied to an August 26 preview, that arithmetic lands in early 2027.

Model Date At Apsara?
Qwen2.5 September 19, 2024 Yes, Apsara 2024
Qwen3 April 29, 2025 No
Qwen3-Next (architecture preview) September 2025 Same week as Apsara 2025
Qwen3-Max September 24, 2025 Yes, Apsara 2025
Qwen3.5 (family built on Qwen3-Next) February 17, 2026 No
Qwen3.8-Max August 3, 2026 No
Qwen3.8-Flash-Next (Qwen 4 architecture preview) August 26, 2026 No
Table 5Where Qwen’s recent launches actually happened

Our read, labeled as our read: Apsara is the most likely place to hear a Qwen 4 date, and a less likely place to download a Qwen 4 model. A Qwen3.8 refresh on the conference stage, with a Qwen 4 timeline in the keynote, fits every pattern above. A full Qwen 4 family on September 22 would be the fastest preview-to-family turn the team has ever done. We would be glad to be wrong; the page flips either way.

§ 06What shipped instead, and what it costs

The reference point for any Qwen 4 claim is the generation Alibaba is actually shipping, and August was the busiest month the Qwen team has had.

Model Date What it is
Qwen3.8-Max August 3 Hosted flagship; “most capable model in the Qwen family to date”
Qwen3.8-2.4T-A95B August 12 Open weights of the flagship base, 2.4T total, 95B active
Qwen3.8-27B August 14 Open dense model
Qwen3.8-Flash-Next August 26 Open Qwen 4 architecture preview, 125B/6B
Qwen3.8-Flash August 28 Production twin on QwenCloud, $0.16 / $0.47 per million tokens, 1M context
Qwen3.8-Max-0902 September 2 Post-trained flagship snapshot, on Code Arena
Table 6Qwen3.8 releases, August and September 2026

Six releases in thirty days, two of them open at flagship scale. That cadence is the second reason to read the Qwen 4 rumor calmly: a team shipping this fast does not need Qwen 4 in September to have news at Apsara. It also explains the “within three months of its prior version” line in Alibaba’s earnings release. For how the current flagship compares with its closest open rival, see GLM-5.3 vs Qwen3.8-Max.

We do not route to Qwen today. CellCog’s tiers run on Claude Fable 5.1, Claude Opus 5 and Gemini 3.8 Flash, and we moved to each on the day it shipped. Qwen matters to us for a different reason, and it is the reason this page exists: the Qwen 4 architecture preview is open, Qwen3.8 opened a Max-class model for the first time, and if the Qwen 4 flagship follows, it becomes the first frontier-class model an AI employee platform could credibly run without a vendor API. That is a decision about the engine underneath every employee we operate, and we would rather have the facts dated before the day comes.

§ 07What we are watching for

We track open-weight frontier models as an operator weighing engines, not as spectators.

  • A date from Alibaba. The Apsara keynote on September 22 is the first scheduled chance. A Qwen team post, an Alibaba Cloud release or a model-card commit would count the same.
  • A Qwen 4 model card. The moment a Qwen/Qwen4-* repository appears on Hugging Face or ModelScope, or QwenCloud’s release notes list a Qwen 4 identifier, this stops being a rumor page.
  • The license line. Qwen3.8-Flash-Next shipped under the Qwen Community License. Whether the Qwen 4 flagship is open, open with restrictions, or API-only is the single fact that changes what this model means for self-hosting.
  • The shape at scale. Does the flagship keep the 125B/6B-with-n-gram-table shape and grow it, or pair the sparse-attention stack with a dense Max-class core? The preview card cannot answer that; the first Qwen 4 card will.
  • Independent replication. Every Flash-Next number is the Qwen team’s. Third-party runs on the preview, and later on Qwen 4, are what make a real comparison possible.
  • Our own routing. If a Qwen 4 model is real, open and better on our axes, we test it on our infrastructure and this page records the result. If it is not, this page records that too.

§ 08The tracker

This page updates when facts change, not when rumors get louder. As of September 4, 2026: the Qwen 4 architecture is public and running in every major inference engine; the Qwen 4 model is not announced, dated, sized or licensed. The most credible window is not the September conference but the months after it, on the team’s own preview-to-family precedent. When Alibaba moves, the confirmed facts and a graded scorecard of the claims above land here the same day.

§ 09Update log

This is a living page; when the story moves, the update lands here.

As of September 4, 2026: page opened. No official movement to log.

Frequently asked6 questions

Q1Has Alibaba announced a Qwen 4 release date?

No. The only official Qwen 4 statements are in the Qwen3.8-Flash-Next release materials of August 26, 2026, which describe an architecture preview and say the full Qwen 4 family will be built on it. Alibaba’s fiscal first-quarter results on August 20 discussed Qwen3.8-Max (‘within three months of its prior version’) and the opened 2.4-trillion-parameter weights, and did not mention Qwen 4. The Apsara Conference page lists September 22 to 24, 2026 and a theme, ‘Intelligence Goes Beyond’, with no model named.

Q2What is Qwen3.8-Flash-Next and how is it related to Qwen 4?

An open-weight multimodal mixture-of-experts model released August 26, 2026, which the Qwen team calls an experimental preview of the architecture that will underpin Qwen 4. Its components: a hybrid of Gated DeltaNet layers and Qwen Sparse Attention, a Gated Residual stream, an N-gram Embedding table of 51B parameters that sits outside the per-token compute budget, and the Muon optimizer. It has 125B parameters with 6B active, plus the 51B table and a 4B multi-token-prediction head, a 262,144-token native context extensible to 1,000,000, and a training cost the team puts at about one ninth of Qwen3.7-Plus. It is not Qwen 4; the product name stays in the 3.8 line. Our full record of that release is on its own page.

Q3Why does open-source code already say qwen4?

Because the runtimes implemented the preview under the name the Qwen team gave the architecture. llama.cpp’s pull request 27742, ‘model: add Qwen3.8-Flash-Next (qwen4exp)’, was opened August 26 and merged August 27. Hugging Face Transformers merged ‘Simplify and fix qwen4 tests’ on August 26 and a ‘[Qwen4 Exp]’ fix on September 1. vLLM has open pull requests titled ‘[Kernel][Qwen] Add qwen4 fuse op’ (August 26), ‘[Qwen4] Support UVA PLE-offload and N-gram parallelism’ (August 29) and ‘[Model][Qwen4] Use ragged PLE IDs for CPU offload’ (August 30). The identifiers mark the architecture generation, not a released model.

Q4Is Qwen 4 Coder 32B-A3B real?

No. An entry by that name, with a June 2026 release date, an Apache 2.0 license and an 82% SWE-bench Verified score, appeared in at least one model database this summer and was picked up by comparison sites. There is no such repository on the Qwen organizations on Hugging Face, ModelScope or GitHub, the claimed score would have beaten every open-weight coding model on record, and the database that carried it says it ingested the entry during a July data incident and removed it in an August reconciliation against official sources. If you see ‘Qwen 4 Coder’ benchmarks, that is where they came from.

Q5Was the anonymous Ox Alpha model on OpenRouter an early Qwen 4?

No. Ox Alpha appeared on OpenRouter’s stealth endpoint on August 20 with a 1M context and multimodal input, and some observers guessed Qwen. Tokenizer and serving fingerprints pointed elsewhere, and Z.ai revealed it on August 26 as GLM-5.3-Flash. As of September 4 no anonymous Arena or OpenRouter checkpoint has been credibly tied to Qwen 4; the Qwen entry added to Code Arena on September 2 is Qwen3.8-Max-0902, a post-trained snapshot of the current flagship.

Q6Should I wait for Qwen 4 before building on Qwen?

No, and the preview is the reason. Anything you run on Qwen3.8-Flash-Next or Qwen3.8-Flash already sits on the Qwen 4 architecture, so the family shipping on top of it is an upgrade to the same stack, not a migration. For our part, CellCog does not route to Qwen today; our tiers run on Claude Fable 5.1, Claude Opus 5 and Gemini 3.8 Flash, and we say so plainly. What a Qwen 4 release changes for an AI employee platform is the open-weight question: a Max-class open model on this architecture would be the first credible self-hosted option at the frontier, and we track that here because the choice of engine sits underneath every employee we run. A complete block of real work on CellCog runs about $25, and the cost depends purely on how much work you assign.

Published 04 September 2026 All Choosing a platform →