# Cognition SWE-2: Benchmarks, the 64% Cost Claim, and the Row It Loses

> Cognition released SWE-2 on Sept 10, 2026: 50.0% FrontierCode, 92.8% Terminal-Bench 2.1, 64% cheaper than Fable 5.1 by its own count. The table, the row it loses.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-09-10
- Canonical (HTML): https://cellcog.ai/blog/cognition-swe-2/
- Section: Guides / Choosing a platform
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- Cognition released SWE-2 on September 10, 2026, calling it 'our most advanced coding model yet' and claiming a score 'within one point of Fable 5.1 while being 64% cheaper' on FrontierCode 1.1 Main. The X headline rounds the savings to 'up to 70% lower cost'.
- The scores in Cognition's table: 50.0% on FrontierCode 1.1 Main (Fable 5.1 50.9%, GPT-6 Astra 53.3%), 73.0% on DeepSWE 1.1 (Fable 5.1 67.4%, Astra 74.1%) and 92.8% on Terminal-Bench 2.1, the only row where SWE-2 beats every model listed.
- The fourth row of the same table is Terminal-Bench 4, where SWE-2 scores 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. 'On par with the frontier' holds on three benchmarks and not on the fourth.
- SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open-weight model that Cognition says had already been through extensive RL for agentic coding. Cognition's RL adds 5 to 6 points on many benchmarks, by its own account.
- It is Cognition's first model with reasoning effort levels, all trained in a single RL run with a cost penalty per level. SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode while scoring higher: 53 mean steps per task against 127.
- Availability is Devin Desktop and CLI today, with Devin Web and Fusion rolling out. Cognition published no API price per token, no context window and no model card. The only dollar figures are per-task benchmark costs for the rival models.
- Every rival number is Cognition's own run: FrontierCode is Cognition's benchmark, and its methodology note says each model was scored in its own harness (Claude Code, Codex, Grok Build, Devin CLI) at its best effort setting, at list prices. Independent replication does not exist yet.

## At a glance

- **What did Cognition release?** SWE-2, its new coding model, on September 10, 2026. Cognition calls it 'our most advanced coding model yet' and 'our closest model yet to the frontier.' It is post-trained from Kimi K3 and is the first SWE model with reasoning effort levels.
- **How good is it?** By Cognition's table: 50.0% on FrontierCode 1.1 Main against 50.9% for Fable 5.1 and 53.3% for GPT-6 Astra; 73.0% on DeepSWE 1.1; 92.8% on Terminal-Bench 2.1, the best in the table; 27.3% on Terminal-Bench 4, roughly half of Fable 5.1 and Astra.
- **What does it cost?** Cognition gives no per-token price. It says SWE-2 is 64% cheaper than Fable 5.1 at the FrontierCode comparison point and about a quarter of GPT-6 Astra's cost, with Fable 5.1 Medium at $3.28 per FrontierCode task as the reference. The X post says 'up to 70% lower cost'.
- **Where can I use it?** Devin Desktop and Devin CLI from September 10, with Devin Web and Fusion rolling out. No standalone API, no model card and no context window have been published as of that day.

Cognition released SWE-2 on September 10, 2026. Its [announcement post](https://cognition.com/blog/swe-2) opens with the claim the whole launch is built on: SWE-2 reaches 50.0% on FrontierCode 1.1 Main, 'within one point of Fable 5.1 while being 64% cheaper'. The [X thread](https://x.com/cognition/status/2098069235733823965), posted at 11:21 AM ET, rounds it up: 'on par with recent frontier models' at 'up to 70% lower cost'.

This page is the launch record: Cognition's full benchmark table as published, the cost claim traced to the only dollar figures in the post, what changed in the model, how the rival numbers were produced, and what is still missing (a price per token, an API, a context window, a model card). Every number here is Cognition's own; no independent evaluation exists yet, and we say so wherever it matters.

## The table, as published

Cognition's post carries one benchmark table. Here it is, with the two frontier columns the X headline is about beside the three others Cognition also ran.

*Table: Cognition's coding benchmark results for SWE-2, September 10, 2026 (Cognition's own runs)*

| Benchmark | SWE-2 | Kimi K3 | Grok 4.6 | Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra | SWE-1.7 |
|---|---|---|---|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 44.2% | 48.0% | 50.9% | 47.5% | 53.3% | 42.0% |
| DeepSWE 1.1 | 73.0% | 68.5% | 67.5% | 67.4% | 72.7% | 74.1% | 37.7% |
| Terminal-Bench 2.1 | 92.8% | 88.3% | 88.4% | 91.4% | 88.8% | 89.9% | 81.5% |
| Terminal-Bench 4 | 27.3% | 21.5% | 20.3% | 55.8% | 37.3% | 57.9% | 7.6% |

Read row by row, the pitch holds on three lines. On FrontierCode 1.1 Main, SWE-2 is 0.9 points behind Fable 5.1 and 3.3 behind GPT-6 Astra. On DeepSWE 1.1 it is ahead of Fable 5.1 by 5.6 points and behind Astra by 1.1. On Terminal-Bench 2.1 it leads everything in the table, including Astra by 2.9 points. Then the fourth line: on Terminal-Bench 4, SWE-2 scores less than half of Fable 5.1 and Astra. Cognition does not hide the row; it is in the table. It is also not in the headline, and it is the reason the claim should be read as 'on par on the benchmarks the launch leads with', which is a narrower and more defensible sentence.

The other thing to hold in view: FrontierCode is Cognition's own benchmark, launched July 17 with a [public leaderboard](https://cognition.com/frontiercode) whose changelog for September 10 reads 'Added SWE-2'. That is normal for a lab post and it is also why the methodology section below matters.

*Two more pictures from Cognition's post, redrawn*

![Bar chart of mean steps per task on FrontierCode 1.1 Main: SWE-1.7 at 127, SWE-2 medium at 53, SWE-2 high at 80, SWE-2 max at 98, with a callout that SWE-2 medium takes 58 percent fewer turns and costs 81 percent less than SWE-1.7](https://cellcog.ai/blog/media/cognition-swe-2/slide-steps.webp)
*Mean steps per task, FrontierCode 1.1 Main: SWE-1.7 127, SWE-2 medium 53, high 80, max 98*

![Four workbenches labeled Claude Code, Codex, Grok Build and Devin CLI, each with a model card reading Fable 5.1, GPT-6 Astra, Grok 4.6 and Kimi K3](https://cellcog.ai/blog/media/cognition-swe-2/slide-harness.webp)
*Cognition's methodology: each rival scored in its own harness, best effort, list prices*

## The cost claim, traced

Cognition publishes no price per million tokens for SWE-2 and no per-task cost for SWE-2 itself. The savings are stated relative to the rivals, and the only dollar figures in the post are the rivals' per-task costs on the two agentic benchmarks. This is what the '64%' is anchored to.

*Table: The dollar figures in Cognition's post (cost per task, list pricing including public discounts, per Cognition)*

| Model and effort | Benchmark | Score | Cost per task |
|---|---|---|---|
| Fable 5.1 Medium | FrontierCode 1.1 Main | 50.9% | $3.28 |
| Fable 5.1 Max | FrontierCode 1.1 Main | 50.3% | $12.83 |
| Fable 5 xhigh | DeepSWE 1.1 | 69.9% | $13.41 |
| Fable 5 Max | DeepSWE 1.1 | 69.7% | $21.63 |
| SWE-2 (any effort) | either | see table above | not published |

Three statements from the post, in its words: SWE-2 is 'within one point of Fable 5.1 while being 64% cheaper' on FrontierCode 1.1 Main; it 'comes within a few points of GPT-6 Astra at a quarter of the cost'; and 'costs assume list pricing, including public discounts'. The X thread's 'up to 70% lower cost' sits between the 64% figure against Fable 5.1 and the roughly 75% implied by 'a quarter of the cost' against Astra. If SWE-2 medium is 64% cheaper than Fable 5.1 Medium at $3.28, the implied figure is about $1.18 per FrontierCode task; that is our arithmetic on Cognition's two numbers, not a number Cognition printed.

What would settle it is a per-token price, which does not exist because SWE-2 is not sold per token. It ships inside Devin, where the cost reaches you as plan usage.

## What changed inside the model

Two things, by Cognition's account. The base moved: SWE-1.7 and SWE-2 share a training recipe, but SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open-weight model that Cognition describes as already heavily RL-trained for agentic coding. Cognition says its own RL still adds 5 to 6 points on many benchmarks on top of that base.

And the model learned to spend less. SWE-2 is the first SWE model with reasoning effort levels, and the post's technical core is how they were trained: a linear cost penalty per effort level, applied in a single RL run, with each penalty tuned to the local slope of the base model's cost-versus-score curve. The practical result Cognition reports is fewer steps for the same or better outcome.

*Table: SWE-1.7 vs SWE-2 on FrontierCode 1.1 Main: mean steps per task (three runs per task, per Cognition)*

| Model and effort | Mean steps per task |
|---|---|
| SWE-1.7 | 127 |
| SWE-2 medium | 53 |
| SWE-2 high | 80 |
| SWE-2 max | 98 |

Cognition's summary of that chart: SWE-2 medium 'scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average'. The rest of the post is infrastructure: a length-weighted reward baseline, NVFP4 and FP8 kernels with quantization-aware training, three times as many RL environments, and verifiers hardened against reward hacking because, in Cognition's words, Kimi K3 is 'a more resourceful model'.

## How the rival numbers were produced

The methodology appendix is short and worth quoting in full: 'For each model-benchmark pair, we report the publicly available result where one exists. Otherwise, we evaluate the model on our internal evaluation framework using the harness for which it was primarily developed: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For each model, we report the best score across reasoning-effort settings.'

Three consequences. First, the Fable 5.1 and Astra columns are Cognition's runs unless a public result existed, and the post does not say which cells are which. Second, best-of-effort-levels is generous to every model, including SWE-2, so the comparison is at least symmetric. Third, and this is the part we agree with most, Cognition scored each model inside the agent product it was built for, because a coding model's score is never the model alone; it is the model plus the loop around it.

## Where you can use it, and where you cannot yet

Cognition's post names four surfaces: 'SWE-2 is available starting today in Devin Desktop and CLI. We're also rolling it out on Devin Web and Fusion.' That is the whole availability statement. There is no standalone API, no model identifier for third-party routers, no context window and no model card in the announcement. Any SWE-2 pricing you see quoted per million tokens on September 10 is someone's estimate.

The company itself had a loud week: a Series E of more than $2 billion at a $48 billion valuation on September 8, a factoring-RSA-260 research post on September 9, and the Dioxus team joining on September 10, the same day as SWE-2. The model is the third item in three days from a company that now has the capital to keep training on multi-trillion-parameter bases.

## Where CellCog stands

We do not route to SWE-2. Our Core and Max tiers run on Claude Fable 5.1 (we moved them on its release day, September 1), Creative runs on Claude Opus 5 and Flash runs on Gemini 3.8 Flash. A Devin-only model is not a routing candidate for us today, and if Cognition ships it as an API we will read the price and the context window before the benchmark table. What we take from this launch is Cognition's own methodology note: the score belongs to the model and its harness together. That is the layer we build.

## What we are watching for

- A per-token price, an API identifier or a model card from Cognition; any of the three turns the cost claim into a number.
- An independent run of SWE-2 on any of the four benchmarks, or on one Cognition did not pick.
- SWE-2 in Devin Web and Fusion (Cognition says rolling out) and in any surface outside Devin.
- Whether the Terminal-Bench 4 gap narrows in a later SWE-2 checkpoint, or whether Cognition addresses it.

## Sources

- Cognition, [Introducing SWE-2: Pushing the Pareto Frontier](https://cognition.com/blog/swe-2), September 10, 2026 (benchmark table, cost statements, availability, methodology appendix).
- Cognition on X, [launch thread](https://x.com/cognition/status/2098069235733823965), September 10, 2026, 11:21 AM ET.
- Cognition, [FrontierCode leaderboard](https://cognition.com/frontiercode), changelog entry September 10, 2026.
- Cognition, [Do it all with Devin: Announcing our Series E](https://cognition.com/blog), September 8, 2026.
- Our [Fable 5.1 record](https://cellcog.ai/blog/fable-5-1-release-date/) and [GPT-6 Astra record](https://cellcog.ai/blog/openai-astra-release-date/) for the two frontier columns' own launch numbers.

## FAQ

**What is SWE-2?**

SWE-2 is Cognition's coding model released September 10, 2026, the successor to SWE-1.7. It is post-trained from Kimi K3, a 2.8-trillion-parameter model, using Cognition's reinforcement-learning recipe scaled for the first time to what the company calls the multi-trillion-parameter regime. It is the first SWE model with reasoning effort levels (medium, high and max), all trained in one RL run.

**Is SWE-2 really on par with Fable 5.1 and GPT-6 Astra?**

On three of the four benchmarks in Cognition's own table, close: 50.0% vs 50.9% and 53.3% on FrontierCode 1.1 Main, 73.0% vs 67.4% and 74.1% on DeepSWE 1.1, and 92.8% vs 91.4% and 89.9% on Terminal-Bench 2.1. On the fourth, Terminal-Bench 4, SWE-2 scores 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. All numbers are Cognition's runs; there is no independent replication yet.

**How much cheaper is SWE-2?**

Cognition's blog says 64% cheaper than Fable 5.1 at the FrontierCode 1.1 Main comparison point and about a quarter of GPT-6 Astra's cost; its X post says 'up to 70% lower cost'. The anchor it publishes is Fable 5.1 Medium at $3.28 per task (Fable 5.1 Max at $12.83). SWE-2's own per-task or per-token price is not published. Cognition says all costs assume list pricing including public discounts.

**Can I call SWE-2 through an API?**

Not as of September 10, 2026. Cognition's announcement names Devin Desktop and CLI as available today and Devin Web and Fusion as rolling out. There is no API identifier, no per-token price, no context window and no model card in the announcement. This page updates when any of those appears.

**Does CellCog run on SWE-2?**

No. CellCog's Core and Max tiers run on Claude Fable 5.1, our Creative mode runs on Claude Opus 5, and Flash runs on Gemini 3.8 Flash. We swap engines under the platform on release days when a model earns it; SWE-2 is available only inside Devin today, so the question does not arise yet. What we do share with Cognition's post is its methodology note: each model was scored inside its own harness, because the harness is a large part of the result.

## Related

- [Fable 5.1 Is Out: Pricing, Benchmarks, and What Actually Changed](https://cellcog.ai/blog/fable-5-1-release-date/index.md)
- [GPT-6 Astra Is Out: Price, Specs, Benchmarks, Rollout, and What Is Still Open](https://cellcog.ai/blog/openai-astra-release-date/index.md)
- [GLM-5.3-Flash Is Ox Alpha: The Reveal, the Specs, and the Real Pricing](https://cellcog.ai/blog/glm-5-3-flash/index.md)
- [Best AI Agent Harnesses: September 2026 Rankings Across the Full Agent Stack](https://cellcog.ai/blog/best-ai-agent-harnesses/index.md)

---

Markdown alternate of https://cellcog.ai/blog/cognition-swe-2/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
