Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact

Cognition SWE-2: Benchmarks, the 64% Cost Claim, and the Row It Loses

At a glanceQuick answers
What did Cognition release?
SWE-2, its new coding model, on September 10, 2026. Cognition calls it ‘our most advanced coding model yet’ and ‘our closest model yet to the frontier.’ It is post-trained from Kimi K3 and is the first SWE model with reasoning effort levels.
How good is it?
By Cognition’s table: 50.0% on FrontierCode 1.1 Main against 50.9% for Fable 5.1 and 53.3% for GPT-6 Astra; 73.0% on DeepSWE 1.1; 92.8% on Terminal-Bench 2.1, the best in the table; 27.3% on Terminal-Bench 4, roughly half of Fable 5.1 and Astra.
What does it cost?
Cognition gives no per-token price. It says SWE-2 is 64% cheaper than Fable 5.1 at the FrontierCode comparison point and about a quarter of GPT-6 Astra’s cost, with Fable 5.1 Medium at $3.28 per FrontierCode task as the reference. The X post says ‘up to 70% lower cost’.
Where can I use it?
Devin Desktop and Devin CLI from September 10, with Devin Web and Fusion rolling out. No standalone API, no model card and no context window have been published as of that day.
Editorial infographic titled SWE-2 vs the frontier: a three-column scoreboard of SWE-2, Fable 5.1 and GPT-6 Astra on FrontierCode 1.1 Main, DeepSWE 1.1 and Terminal-Bench 2.1, a teal badge reading 64% cheaper than Fable 5.1, and a mustard strip with the Terminal-Bench 4 scores
Fig 0Cognition's numbers, redrawn. The scoreboard is the launch pitch; the mustard strip is the fourth row of the same table.

Cognition released SWE-2 on September 10, 2026. Its announcement post opens with the claim the whole launch is built on: SWE-2 reaches 50.0% on FrontierCode 1.1 Main, ‘within one point of Fable 5.1 while being 64% cheaper’. The X thread, posted at 11:21 AM ET, rounds it up: ‘on par with recent frontier models’ at ‘up to 70% lower cost’.

This page is the launch record: Cognition’s full benchmark table as published, the cost claim traced to the only dollar figures in the post, what changed in the model, how the rival numbers were produced, and what is still missing (a price per token, an API, a context window, a model card). Every number here is Cognition’s own; no independent evaluation exists yet, and we say so wherever it matters.

On this page · 8 sectionsOpen
  1. The table, as published
  2. The cost claim, traced
  3. What changed inside the model
  4. How the rival numbers were produced
  5. Where you can use it, and where you cannot yet
  6. Where CellCog stands
  7. What we are watching for
  8. Sources
Key points7 · 10 min full read
  1. A pair of teal scissors cutting a paper price tag, the cut corner falling away: the launch pitch is the cost, not the score.
    Cognition released SWE-2 on September 10, 2026, calling it ‘our most advanced coding model yet’ and claiming a score ‘within one point of Fable 5.1 while being 64% cheaper’ on FrontierCode 1.1 Main. The X headline rounds the savings to ‘up to 70% lower cost’.
  2. Five bars with the coral one a hair taller than the rest and a small flag on top: one row where SWE-2 comes first.
    The scores in Cognition’s table: 50.0% on FrontierCode 1.1 Main (Fable 5.1 50.9%, GPT-6 Astra 53.3%), 73.0% on DeepSWE 1.1 (Fable 5.1 67.4%, Astra 74.1%) and 92.8% on Terminal-Bench 2.1, the only row where SWE-2 beats every model listed.
  3. A track hurdle knocked over on its side beside one still standing: the fourth benchmark row.
    The fourth row of the same table is Terminal-Bench 4, where SWE-2 scores 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. ‘On par with the frontier’ holds on three benchmarks and not on the fourth.
  4. A young sapling with teal leaves growing from the top of a cut tree stump: a new model grown on Kimi K3.
    SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open-weight model that Cognition says had already been through extensive RL for agentic coding. Cognition’s RL adds 5 to 6 points on many benchmarks, by its own account.
  5. A round control dial with three notches, the pointer on the middle one glowing coral: the medium effort level.
    It is Cognition’s first model with reasoning effort levels, all trained in a single RL run with a cost penalty per level. SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode while scoring higher: 53 mean steps per task against 127.
  6. A single closed teal door with one coral key hanging on a hook beside it: SWE-2 opens inside Devin only.
    Availability is Devin Desktop and CLI today, with Devin Web and Fusion rolling out. Cognition published no API price per token, no context window and no model card. The only dollar figures are per-task benchmark costs for the rival models.
  7. A magnifying glass held over a wooden ruler with one coral tick: the measurement is the vendor's own.
    Every rival number is Cognition’s own run: FrontierCode is Cognition’s benchmark, and its methodology note says each model was scored in its own harness (Claude Code, Codex, Grok Build, Devin CLI) at its best effort setting, at list prices. Independent replication does not exist yet.

§ 01The table, as published

Cognition’s post carries one benchmark table. Here it is, with the two frontier columns the X headline is about beside the three others Cognition also ran.

Benchmark SWE-2 Kimi K3 Grok 4.6 Fable 5.1 GPT-5.6 Sol GPT-6 Astra SWE-1.7
FrontierCode 1.1 Main 50.0% 44.2% 48.0% 50.9% 47.5% 53.3% 42.0%
DeepSWE 1.1 73.0% 68.5% 67.5% 67.4% 72.7% 74.1% 37.7%
Terminal-Bench 2.1 92.8% 88.3% 88.4% 91.4% 88.8% 89.9% 81.5%
Terminal-Bench 4 27.3% 21.5% 20.3% 55.8% 37.3% 57.9% 7.6%
Scroll to compare all columns
Table 1Cognition’s coding benchmark results for SWE-2, September 10, 2026 (Cognition’s own runs)
Terminal-Bench 4 in Cognition's own table: the row SWE-2 losesBar chart of Terminal-Bench 4 scores from Cognition's table: GPT-6 Astra 57.9, Fable 5.1 55.8, GPT-5.6 Sol 37.3, SWE-2 highlighted at 27.3, Kimi K3 21.5, Grok 4.6 20.3, SWE-1.7 7.6GPT-6 Astra57.9Fable 5.155.8GPT-5.6 Sol37.3SWE-227.3Kimi K321.5Grok 4.620.3SWE-1.77.6Terminal-Bench 4 in Cognition's own table: the row SWE-2 losesBar chart of Terminal-Bench 4 scores from Cognition's table: GPT-6 Astra 57.9, Fable 5.1 55.8, GPT-5.6 Sol 37.3, SWE-2 highlighted at 27.3, Kimi K3 21.5, Grok 4.6 20.3, SWE-1.7 7.6GPT-6 Astra57.9Fable 5.155.8GPT-5.6 Sol37.3SWE-227.3Kimi K321.5Grok 4.620.3SWE-1.77.6
Fig 1Terminal-Bench 4 in Cognition's own table: the row SWE-2 loses

Read row by row, the pitch holds on three lines. On FrontierCode 1.1 Main, SWE-2 is 0.9 points behind Fable 5.1 and 3.3 behind GPT-6 Astra. On DeepSWE 1.1 it is ahead of Fable 5.1 by 5.6 points and behind Astra by 1.1. On Terminal-Bench 2.1 it leads everything in the table, including Astra by 2.9 points. Then the fourth line: on Terminal-Bench 4, SWE-2 scores less than half of Fable 5.1 and Astra. Cognition does not hide the row; it is in the table. It is also not in the headline, and it is the reason the claim should be read as ‘on par on the benchmarks the launch leads with’, which is a narrower and more defensible sentence.

The other thing to hold in view: FrontierCode is Cognition’s own benchmark, launched July 17 with a public leaderboard whose changelog for September 10 reads ‘Added SWE-2’. That is normal for a lab post and it is also why the methodology section below matters.

Bar chart of mean steps per task on FrontierCode 1.1 Main: SWE-1.7 at 127, SWE-2 medium at 53, SWE-2 high at 80, SWE-2 max at 98, with a callout that SWE-2 medium takes 58 percent fewer turns and costs 81 percent less than SWE-1.7

01Mean steps per task, FrontierCode 1.1 Main: SWE-1.7 127, SWE-2 medium 53, high 80, max 98

Four workbenches labeled Claude Code, Codex, Grok Build and Devin CLI, each with a model card reading Fable 5.1, GPT-6 Astra, Grok 4.6 and Kimi K3

02Cognition's methodology: each rival scored in its own harness, best effort, list prices

1 / 2
Fig 2Two more pictures from Cognition's post, redrawn

§ 02The cost claim, traced

Cognition publishes no price per million tokens for SWE-2 and no per-task cost for SWE-2 itself. The savings are stated relative to the rivals, and the only dollar figures in the post are the rivals’ per-task costs on the two agentic benchmarks. This is what the ‘64%’ is anchored to.

Model and effort Benchmark Score Cost per task
Fable 5.1 Medium FrontierCode 1.1 Main 50.9% $3.28
Fable 5.1 Max FrontierCode 1.1 Main 50.3% $12.83
Fable 5 xhigh DeepSWE 1.1 69.9% $13.41
Fable 5 Max DeepSWE 1.1 69.7% $21.63
SWE-2 (any effort) either see table above not published
Table 2The dollar figures in Cognition’s post (cost per task, list pricing including public discounts, per Cognition)

Three statements from the post, in its words: SWE-2 is ‘within one point of Fable 5.1 while being 64% cheaper’ on FrontierCode 1.1 Main; it ‘comes within a few points of GPT-6 Astra at a quarter of the cost’; and ‘costs assume list pricing, including public discounts’. The X thread’s ‘up to 70% lower cost’ sits between the 64% figure against Fable 5.1 and the roughly 75% implied by ‘a quarter of the cost’ against Astra. If SWE-2 medium is 64% cheaper than Fable 5.1 Medium at $3.28, the implied figure is about $1.18 per FrontierCode task; that is our arithmetic on Cognition’s two numbers, not a number Cognition printed.

What would settle it is a per-token price, which does not exist because SWE-2 is not sold per token. It ships inside Devin, where the cost reaches you as plan usage.

§ 03What changed inside the model

Two things, by Cognition’s account. The base moved: SWE-1.7 and SWE-2 share a training recipe, but SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open-weight model that Cognition describes as already heavily RL-trained for agentic coding. Cognition says its own RL still adds 5 to 6 points on many benchmarks on top of that base.

And the model learned to spend less. SWE-2 is the first SWE model with reasoning effort levels, and the post’s technical core is how they were trained: a linear cost penalty per effort level, applied in a single RL run, with each penalty tuned to the local slope of the base model’s cost-versus-score curve. The practical result Cognition reports is fewer steps for the same or better outcome.

Model and effort Mean steps per task
SWE-1.7 127
SWE-2 medium 53
SWE-2 high 80
SWE-2 max 98
Table 3SWE-1.7 vs SWE-2 on FrontierCode 1.1 Main: mean steps per task (three runs per task, per Cognition)
Mean steps per task on FrontierCode 1.1 Main, per CognitionBar chart of mean steps per task: SWE-1.7 at 127, SWE-2 medium highlighted at 53, SWE-2 high at 80, SWE-2 max at 98SWE-1.7127SWE-2 medium53SWE-2 high80SWE-2 max98Mean steps per task on FrontierCode 1.1 Main, per CognitionBar chart of mean steps per task: SWE-1.7 at 127, SWE-2 medium highlighted at 53, SWE-2 high at 80, SWE-2 max at 98SWE-1.7127SWE-2 medium53SWE-2 high80SWE-2 max98
Fig 3Mean steps per task on FrontierCode 1.1 Main, per Cognition

Cognition’s summary of that chart: SWE-2 medium ‘scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average’. The rest of the post is infrastructure: a length-weighted reward baseline, NVFP4 and FP8 kernels with quantization-aware training, three times as many RL environments, and verifiers hardened against reward hacking because, in Cognition’s words, Kimi K3 is ‘a more resourceful model’.

§ 04How the rival numbers were produced

The methodology appendix is short and worth quoting in full: ‘For each model-benchmark pair, we report the publicly available result where one exists. Otherwise, we evaluate the model on our internal evaluation framework using the harness for which it was primarily developed: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For each model, we report the best score across reasoning-effort settings.’

Three consequences. First, the Fable 5.1 and Astra columns are Cognition’s runs unless a public result existed, and the post does not say which cells are which. Second, best-of-effort-levels is generous to every model, including SWE-2, so the comparison is at least symmetric. Third, and this is the part we agree with most, Cognition scored each model inside the agent product it was built for, because a coding model’s score is never the model alone; it is the model plus the loop around it.

§ 05Where you can use it, and where you cannot yet

Cognition’s post names four surfaces: ‘SWE-2 is available starting today in Devin Desktop and CLI. We’re also rolling it out on Devin Web and Fusion.’ That is the whole availability statement. There is no standalone API, no model identifier for third-party routers, no context window and no model card in the announcement. Any SWE-2 pricing you see quoted per million tokens on September 10 is someone’s estimate.

The company itself had a loud week: a Series E of more than $2 billion at a $48 billion valuation on September 8, a factoring-RSA-260 research post on September 9, and the Dioxus team joining on September 10, the same day as SWE-2. The model is the third item in three days from a company that now has the capital to keep training on multi-trillion-parameter bases.

§ 06Where CellCog stands

We do not route to SWE-2. Our Core and Max tiers run on Claude Fable 5.1 (we moved them on its release day, September 1), Creative runs on Claude Opus 5 and Flash runs on Gemini 3.8 Flash. A Devin-only model is not a routing candidate for us today, and if Cognition ships it as an API we will read the price and the context window before the benchmark table. What we take from this launch is Cognition’s own methodology note: the score belongs to the model and its harness together. That is the layer we build.

§ 07What we are watching for

  • A per-token price, an API identifier or a model card from Cognition; any of the three turns the cost claim into a number.
  • An independent run of SWE-2 on any of the four benchmarks, or on one Cognition did not pick.
  • SWE-2 in Devin Web and Fusion (Cognition says rolling out) and in any surface outside Devin.
  • Whether the Terminal-Bench 4 gap narrows in a later SWE-2 checkpoint, or whether Cognition addresses it.

§ 08Sources

Frequently asked5 questions

Q1What is SWE-2?

SWE-2 is Cognition’s coding model released September 10, 2026, the successor to SWE-1.7. It is post-trained from Kimi K3, a 2.8-trillion-parameter model, using Cognition’s reinforcement-learning recipe scaled for the first time to what the company calls the multi-trillion-parameter regime. It is the first SWE model with reasoning effort levels (medium, high and max), all trained in one RL run.

Q2Is SWE-2 really on par with Fable 5.1 and GPT-6 Astra?

On three of the four benchmarks in Cognition’s own table, close: 50.0% vs 50.9% and 53.3% on FrontierCode 1.1 Main, 73.0% vs 67.4% and 74.1% on DeepSWE 1.1, and 92.8% vs 91.4% and 89.9% on Terminal-Bench 2.1. On the fourth, Terminal-Bench 4, SWE-2 scores 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. All numbers are Cognition’s runs; there is no independent replication yet.

Q3How much cheaper is SWE-2?

Cognition’s blog says 64% cheaper than Fable 5.1 at the FrontierCode 1.1 Main comparison point and about a quarter of GPT-6 Astra’s cost; its X post says ‘up to 70% lower cost’. The anchor it publishes is Fable 5.1 Medium at $3.28 per task (Fable 5.1 Max at $12.83). SWE-2’s own per-task or per-token price is not published. Cognition says all costs assume list pricing including public discounts.

Q4Can I call SWE-2 through an API?

Not as of September 10, 2026. Cognition’s announcement names Devin Desktop and CLI as available today and Devin Web and Fusion as rolling out. There is no API identifier, no per-token price, no context window and no model card in the announcement. This page updates when any of those appears.

Q5Does CellCog run on SWE-2?

No. CellCog’s Core and Max tiers run on Claude Fable 5.1, our Creative mode runs on Claude Opus 5, and Flash runs on Gemini 3.8 Flash. We swap engines under the platform on release days when a model earns it; SWE-2 is available only inside Devin today, so the question does not arise yet. What we do share with Cognition’s post is its methodology note: each model was scored inside its own harness, because the harness is a large part of the result.

Published 10 September 2026 All Choosing a platform →