Cognition released SWE-2 on September 10, 2026. Its announcement post opens with the claim the whole launch is built on: SWE-2 reaches 50.0% on FrontierCode 1.1 Main, ‘within one point of Fable 5.1 while being 64% cheaper’. The X thread, posted at 11:21 AM ET, rounds it up: ‘on par with recent frontier models’ at ‘up to 70% lower cost’.
This page is the launch record: Cognition’s full benchmark table as published, the cost claim traced to the only dollar figures in the post, what changed in the model, how the rival numbers were produced, and what is still missing (a price per token, an API, a context window, a model card). Every number here is Cognition’s own; no independent evaluation exists yet, and we say so wherever it matters.
On this page · 8 sectionsOpen
Cognition released SWE-2 on September 10, 2026, calling it ‘our most advanced coding model yet’ and claiming a score ‘within one point of Fable 5.1 while being 64% cheaper’ on FrontierCode 1.1 Main. The X headline rounds the savings to ‘up to 70% lower cost’.
The scores in Cognition’s table: 50.0% on FrontierCode 1.1 Main (Fable 5.1 50.9%, GPT-6 Astra 53.3%), 73.0% on DeepSWE 1.1 (Fable 5.1 67.4%, Astra 74.1%) and 92.8% on Terminal-Bench 2.1, the only row where SWE-2 beats every model listed.
The fourth row of the same table is Terminal-Bench 4, where SWE-2 scores 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. ‘On par with the frontier’ holds on three benchmarks and not on the fourth.
SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open-weight model that Cognition says had already been through extensive RL for agentic coding. Cognition’s RL adds 5 to 6 points on many benchmarks, by its own account.
It is Cognition’s first model with reasoning effort levels, all trained in a single RL run with a cost penalty per level. SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode while scoring higher: 53 mean steps per task against 127.
Availability is Devin Desktop and CLI today, with Devin Web and Fusion rolling out. Cognition published no API price per token, no context window and no model card. The only dollar figures are per-task benchmark costs for the rival models.
Every rival number is Cognition’s own run: FrontierCode is Cognition’s benchmark, and its methodology note says each model was scored in its own harness (Claude Code, Codex, Grok Build, Devin CLI) at its best effort setting, at list prices. Independent replication does not exist yet.
§ 01The table, as published
Cognition’s post carries one benchmark table. Here it is, with the two frontier columns the X headline is about beside the three others Cognition also ran.
| Benchmark | SWE-2 | Kimi K3 | Grok 4.6 | Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra | SWE-1.7 |
|---|---|---|---|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 44.2% | 48.0% | 50.9% | 47.5% | 53.3% | 42.0% |
| DeepSWE 1.1 | 73.0% | 68.5% | 67.5% | 67.4% | 72.7% | 74.1% | 37.7% |
| Terminal-Bench 2.1 | 92.8% | 88.3% | 88.4% | 91.4% | 88.8% | 89.9% | 81.5% |
| Terminal-Bench 4 | 27.3% | 21.5% | 20.3% | 55.8% | 37.3% | 57.9% | 7.6% |
Read row by row, the pitch holds on three lines. On FrontierCode 1.1 Main, SWE-2 is 0.9 points behind Fable 5.1 and 3.3 behind GPT-6 Astra. On DeepSWE 1.1 it is ahead of Fable 5.1 by 5.6 points and behind Astra by 1.1. On Terminal-Bench 2.1 it leads everything in the table, including Astra by 2.9 points. Then the fourth line: on Terminal-Bench 4, SWE-2 scores less than half of Fable 5.1 and Astra. Cognition does not hide the row; it is in the table. It is also not in the headline, and it is the reason the claim should be read as ‘on par on the benchmarks the launch leads with’, which is a narrower and more defensible sentence.
The other thing to hold in view: FrontierCode is Cognition’s own benchmark, launched July 17 with a public leaderboard whose changelog for September 10 reads ‘Added SWE-2’. That is normal for a lab post and it is also why the methodology section below matters.
§ 02The cost claim, traced
Cognition publishes no price per million tokens for SWE-2 and no per-task cost for SWE-2 itself. The savings are stated relative to the rivals, and the only dollar figures in the post are the rivals’ per-task costs on the two agentic benchmarks. This is what the ‘64%’ is anchored to.
| Model and effort | Benchmark | Score | Cost per task |
|---|---|---|---|
| Fable 5.1 Medium | FrontierCode 1.1 Main | 50.9% | $3.28 |
| Fable 5.1 Max | FrontierCode 1.1 Main | 50.3% | $12.83 |
| Fable 5 xhigh | DeepSWE 1.1 | 69.9% | $13.41 |
| Fable 5 Max | DeepSWE 1.1 | 69.7% | $21.63 |
| SWE-2 (any effort) | either | see table above | not published |
Three statements from the post, in its words: SWE-2 is ‘within one point of Fable 5.1 while being 64% cheaper’ on FrontierCode 1.1 Main; it ‘comes within a few points of GPT-6 Astra at a quarter of the cost’; and ‘costs assume list pricing, including public discounts’. The X thread’s ‘up to 70% lower cost’ sits between the 64% figure against Fable 5.1 and the roughly 75% implied by ‘a quarter of the cost’ against Astra. If SWE-2 medium is 64% cheaper than Fable 5.1 Medium at $3.28, the implied figure is about $1.18 per FrontierCode task; that is our arithmetic on Cognition’s two numbers, not a number Cognition printed.
What would settle it is a per-token price, which does not exist because SWE-2 is not sold per token. It ships inside Devin, where the cost reaches you as plan usage.
§ 03What changed inside the model
Two things, by Cognition’s account. The base moved: SWE-1.7 and SWE-2 share a training recipe, but SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open-weight model that Cognition describes as already heavily RL-trained for agentic coding. Cognition says its own RL still adds 5 to 6 points on many benchmarks on top of that base.
And the model learned to spend less. SWE-2 is the first SWE model with reasoning effort levels, and the post’s technical core is how they were trained: a linear cost penalty per effort level, applied in a single RL run, with each penalty tuned to the local slope of the base model’s cost-versus-score curve. The practical result Cognition reports is fewer steps for the same or better outcome.
| Model and effort | Mean steps per task |
|---|---|
| SWE-1.7 | 127 |
| SWE-2 medium | 53 |
| SWE-2 high | 80 |
| SWE-2 max | 98 |
Cognition’s summary of that chart: SWE-2 medium ‘scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average’. The rest of the post is infrastructure: a length-weighted reward baseline, NVFP4 and FP8 kernels with quantization-aware training, three times as many RL environments, and verifiers hardened against reward hacking because, in Cognition’s words, Kimi K3 is ‘a more resourceful model’.
§ 04How the rival numbers were produced
The methodology appendix is short and worth quoting in full: ‘For each model-benchmark pair, we report the publicly available result where one exists. Otherwise, we evaluate the model on our internal evaluation framework using the harness for which it was primarily developed: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For each model, we report the best score across reasoning-effort settings.’
Three consequences. First, the Fable 5.1 and Astra columns are Cognition’s runs unless a public result existed, and the post does not say which cells are which. Second, best-of-effort-levels is generous to every model, including SWE-2, so the comparison is at least symmetric. Third, and this is the part we agree with most, Cognition scored each model inside the agent product it was built for, because a coding model’s score is never the model alone; it is the model plus the loop around it.
§ 05Where you can use it, and where you cannot yet
Cognition’s post names four surfaces: ‘SWE-2 is available starting today in Devin Desktop and CLI. We’re also rolling it out on Devin Web and Fusion.’ That is the whole availability statement. There is no standalone API, no model identifier for third-party routers, no context window and no model card in the announcement. Any SWE-2 pricing you see quoted per million tokens on September 10 is someone’s estimate.
The company itself had a loud week: a Series E of more than $2 billion at a $48 billion valuation on September 8, a factoring-RSA-260 research post on September 9, and the Dioxus team joining on September 10, the same day as SWE-2. The model is the third item in three days from a company that now has the capital to keep training on multi-trillion-parameter bases.
§ 06Where CellCog stands
We do not route to SWE-2. Our Core and Max tiers run on Claude Fable 5.1 (we moved them on its release day, September 1), Creative runs on Claude Opus 5 and Flash runs on Gemini 3.8 Flash. A Devin-only model is not a routing candidate for us today, and if Cognition ships it as an API we will read the price and the context window before the benchmark table. What we take from this launch is Cognition’s own methodology note: the score belongs to the model and its harness together. That is the layer we build.
§ 07What we are watching for
- A per-token price, an API identifier or a model card from Cognition; any of the three turns the cost claim into a number.
- An independent run of SWE-2 on any of the four benchmarks, or on one Cognition did not pick.
- SWE-2 in Devin Web and Fusion (Cognition says rolling out) and in any surface outside Devin.
- Whether the Terminal-Bench 4 gap narrows in a later SWE-2 checkpoint, or whether Cognition addresses it.
§ 08Sources
- Cognition, Introducing SWE-2: Pushing the Pareto Frontier, September 10, 2026 (benchmark table, cost statements, availability, methodology appendix).
- Cognition on X, launch thread, September 10, 2026, 11:21 AM ET.
- Cognition, FrontierCode leaderboard, changelog entry September 10, 2026.
- Cognition, Do it all with Devin: Announcing our Series E, September 8, 2026.
- Our Fable 5.1 record and GPT-6 Astra record for the two frontier columns’ own launch numbers.
Q1What is SWE-2?
SWE-2 is Cognition’s coding model released September 10, 2026, the successor to SWE-1.7. It is post-trained from Kimi K3, a 2.8-trillion-parameter model, using Cognition’s reinforcement-learning recipe scaled for the first time to what the company calls the multi-trillion-parameter regime. It is the first SWE model with reasoning effort levels (medium, high and max), all trained in one RL run.
Q2Is SWE-2 really on par with Fable 5.1 and GPT-6 Astra?
On three of the four benchmarks in Cognition’s own table, close: 50.0% vs 50.9% and 53.3% on FrontierCode 1.1 Main, 73.0% vs 67.4% and 74.1% on DeepSWE 1.1, and 92.8% vs 91.4% and 89.9% on Terminal-Bench 2.1. On the fourth, Terminal-Bench 4, SWE-2 scores 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. All numbers are Cognition’s runs; there is no independent replication yet.
Q3How much cheaper is SWE-2?
Cognition’s blog says 64% cheaper than Fable 5.1 at the FrontierCode 1.1 Main comparison point and about a quarter of GPT-6 Astra’s cost; its X post says ‘up to 70% lower cost’. The anchor it publishes is Fable 5.1 Medium at $3.28 per task (Fable 5.1 Max at $12.83). SWE-2’s own per-task or per-token price is not published. Cognition says all costs assume list pricing including public discounts.
Q4Can I call SWE-2 through an API?
Not as of September 10, 2026. Cognition’s announcement names Devin Desktop and CLI as available today and Devin Web and Fusion as rolling out. There is no API identifier, no per-token price, no context window and no model card in the announcement. This page updates when any of those appears.
Q5Does CellCog run on SWE-2?
No. CellCog’s Core and Max tiers run on Claude Fable 5.1, our Creative mode runs on Claude Opus 5, and Flash runs on Gemini 3.8 Flash. We swap engines under the platform on release days when a model earns it; SWE-2 is available only inside Devin today, so the question does not arise yet. What we do share with Cognition’s post is its methodology note: each model was scored inside its own harness, because the harness is a large part of the result.


