The week had two frontier launches. On September 1, 2026, Anthropic made Claude Fable 5.1 generally available and OpenAI said Astra was coming soon. On September 3, OpenAI launched GPT-6 Astra in a press briefing and started a phased rollout. “Astra vs Fable 5.1” went from a search phrase to a real question in 48 hours.
By 4 pm ET on launch day the two models were specified to nearly the same degree, and the specs match to the dollar: $10 and $50 per million tokens for each, 1M-class context windows, 128K output. What differs is access, the cache rate, the safety posture, and the evidence: Fable 5.1 is generally available with Anthropic’s coding benchmarks; Astra is reaching Trusted Access enterprises first, with a 117-page system card out at about 5 pm ET, a launch post filed under Safety that evening, and its capability benchmarks quoted by reporters from a release text that is still not on openai.com. So this page compares what is actually knowable as of September 3, dated and sourced, in the same living format as our Fable 5.1 tracker and GPT-6 Astra tracker, and grades its own September 1 predictions at the bottom.
On this page · 11 sectionsOpen
- The comparison, as of September 3, 2026
- The one benchmark with both names on it
- Same price, different bill
- Where the money went last week
- What Fable 5.1 is
- What GPT-6 Astra is
- Why there is still no verdict
- OpenAI’s head-to-head table, and what it is worth
- What each is for, and what to do today
- How this page’s own predictions scored
- Update log
- The week has two frontier launches: Anthropic made Claude Fable 5.1 generally available on September 1, 2026, and OpenAI launched GPT-6 Astra on September 3 through a press briefing, with a phased rollout that started with its Daybreak cybersecurity program.
- Both are now specified, and they match to the dollar: $10 per million input tokens and $50 per million output tokens for each, API ids claude-fable-5-1 and gpt-6-astra, context windows of 1M and 1,050,000 tokens with 128K output. The difference is in the cache: Fable 5.1 reads cached input at $0.25 per million, Astra at $1.
- OpenAI’s product post, published after launch day, puts Claude Fable 5.1 in Astra’s benchmark table on four rows, with Anthropic’s published figures beside OpenAI’s: Terminal-Bench 4.0 57.9% to 55.8%, Terminal-Bench Science 64.6% to 52.6%, AutomationBench 41.4% to 31.4%, BenchCAD 95.9% to 84.3%. On computer use, Agents’ Last Exam 59.3% and OSWorld 72.6%, the Anthropic comparison is Opus 5 (55.5% and 70.2%), not Fable 5.1. OpenAI’s settings throughout.
- Astra’s other launch numbers, all OpenAI-reported: computer use 72.6% at about 40 minutes per task (about 47% less time than GPT-5.6 Sol), 0% out-of-scope actions on a post-Hugging-Face evaluation versus 48.2% for Sol, and the September 1 cyber results. Fable 5.1’s are coding and knowledge work: Terminal-Bench 4.0 55.8%, CursorBench 3.2.0 73.4%.
- The safety postures differ in kind: Anthropic ships Fable 5.1 with production safeguards and a restricted Mythos 5.1 variant; OpenAI designated Astra the first Critical-level cyber model, gates its advanced cyber capabilities to Daybreak, and acknowledges Astra is harder to monitor because it can conceal step-by-step reasoning.
- What to use today: Fable 5.1 if you need general availability now; Astra as OpenAI’s rollout reaches your plan or account, with the developer docs no longer naming a gate as of September 4. CellCog’s Agent Max and Team Max run on Fable 5.1 as of September 1; no platform offers Astra yet, and OpenAI puts Azure and Bedrock at ‘the coming days’.
- This page said it would be rewritten the day Astra shipped, with a scorecard. It was, on this URL, on September 3. The independent benchmarks it promised to add do not exist yet.
§ 01The comparison, as of September 3, 2026
| Claude Fable 5.1 | GPT-6 Astra | |
|---|---|---|
| Status | Generally available since September 1 | Launched September 3; phased rollout under way |
| Release date | September 1, 2026 | September 3, 2026 |
| Price | $10 / $50 per million input / output tokens; cache reads $0.25 per million | $10 / $50 per million input / output tokens; cached input $1; 2x input and 1.5x output above 272K input tokens |
| API model id | claude-fable-5-1 | gpt-6-astra |
| Context window | 1M tokens, 128K output | 1,050,000 tokens, 128K output; April 30, 2026 knowledge cutoff |
| Who can use it today | Anyone via Anthropic’s API and partners; CellCog Agent Max and Team Max | Rolling out: the developer docs no longer name an access gate and publish usage-tier rate limits, Codex CLI ships it as the bundled default, and OpenAI’s product post puts general API, Azure and Bedrock access at “the coming days”; ChatGPT Plus, Pro, Business and Enterprise per the product post (Plus not in Chat per the Help Center); not Cursor |
| Safety posture | Production safeguards; Mythos 5.1 is the same model with permissive safeguards under restricted access | First model designated Critical for cybersecurity; advanced cyber access gated to Daybreak; misalignment monitoring on all external tool use; system card states monitorability decreased vs GPT-5.6 Sol |
| Vendor benchmarks | Terminal-Bench 4.0 55.8%, Terminal-Bench-Science 52.6%, CursorBench 3.2.0 73.4%, AutomationBench 31.4% (Anthropic) | OSWorld 2.0 72.6% at about 40 min per task; Agents’ Last Exam 59.3%; Terminal-Bench 4.0 57.9%; AutomationBench 41.4%; BenchCAD 95.9%; ExploitBench 100% (OpenAI) |
| Head-to-head | No Anthropic number against Astra | OpenAI’s table beside Anthropic’s published Fable 5.1 figures: Astra ahead on Terminal-Bench 4.0 (57.9% vs 55.8%), Terminal-Bench Science (64.6% vs 52.6%), AutomationBench (41.4% vs 31.4%) and BenchCAD (95.9% vs 84.3%); no Fable 5.1 number on Agents’ Last Exam or OSWorld; OpenAI’s September 9 post adds two OpenAI-run figures against Fable 5.1, about 63% lower estimated API cost per task on Terminal-Bench 4.0 and 74.7% fewer unintended outcomes on an internal computer-use safety benchmark |
| Independent tests | None against Astra | None |
| Positioning | Long-running coding and knowledge work; avoids shortcuts, fixes root causes | Computer and browser use, software engineering, professional work, science; “most aligned” |
| What you can do today | Use it | Use it if you are in Daybreak; otherwise wait days |
§ 02The one benchmark with both names on it
OpenAI chose one place to stand Astra next to Claude, and it chose the previous Claude. Per OpenAI’s release as quoted by ZDNet, Astra scores 59.3% on Agent’s Last Exam, an external benchmark of professional-grade agent work, against 48.7% for Claude Fable 5 and 55.5% for Claude Opus 5.
| Model | Score | Who ran it |
|---|---|---|
| GPT-6 Astra | 59.3% | OpenAI, September 3 |
| Claude Opus 5 | 55.5% | OpenAI, September 3 |
| Claude Fable 5 | 48.7% | OpenAI, September 3 |
| Claude Fable 5.1 | Not published | Anthropic has released no score |
Read it for what it is. The gap is real if the runs are fair, and OpenAI has every incentive to pick the benchmark and the comparison model that flatter Astra. Fable 5.1 shipped two days before the briefing and has no number here; Anthropic’s own Fable 5.1 figures are on Terminal-Bench, CursorBench and AutomationBench, where OpenAI reports nothing. TechCrunch adds that OpenAI showed Astra ahead of “Anthropic’s Fable” on bug finding, terminal tasks and codebase questions, again without public figures. Until a third party runs both models on the same tasks, the honest column header is “vendor benchmarks,” and that is what the table says.
§ 03Same price, different bill
OpenAI priced Astra to the dollar against Fable 5.1 on the two numbers everyone quotes, and differently on the one that matters most for agents.
| Claude Fable 5.1 | GPT-6 Astra | |
|---|---|---|
| Input | $10 | $10 (2x above 272K input tokens) |
| Output | $50 | $50 (1.5x above 272K input tokens) |
| Cached input | $0.25 | $1.00 |
| Context window | 1M tokens, 128K output | 1,050,000 tokens, 128K output |
| Discounted lanes | Batch API and prompt caching | Batch and Flex at 50%; Fast mode at 2x |
| Source | Anthropic pricing, September 1 | OpenAI pricing page and model page, September 3 |
A standing agent re-reads its instructions, memory and conversation on every turn, so most of the tokens it pays for are cached input. At Anthropic’s rate that re-read costs $0.25 per million; at OpenAI’s it costs $1. On the headline the two are tied; on the bill, a long-running agent pays more on Astra unless its work is dominated by fresh output, and Astra’s surcharge above 272K input tokens points the same way. These are list prices read on September 3; real-workload comparisons will settle what the gap is in practice.
§ 04Where the money went last week
OpenRouter posted the first market reading on September 15 at 16:28 UTC: “OpenRouter users spent more on OpenAI models than on Anthropic models last week. This hasn’t happened for more than 2.5 years”. The week is September 7 to 13, the first full week with Astra on the platform. OfficeChai, reading OpenRouter’s weekly wallet-share data, puts the last OpenAI lead in the week of February 26, 2024, Astra alone at roughly 19% of spend, and OpenAI’s four-model line just past the 50% mark of the two labs’ combined total. Three cautions before reading it as a verdict. OpenRouter is one router, weighted toward developers who shop models by price and try new ones early, so a launch week flatters the new model. Spend is not usage: at $10 and $50 a million tokens Astra earns its share on fewer tokens than the cheaper GPT-5.6 line does. And one week is one week; the same data had Anthropic at 75 to 80% of combined spend through most of 2024 and 2025. What the number does settle is the question this page was built for: enough teams moved real money onto Astra in its first full week that the pricing table above now describes a live choice, not a launch announcement.
§ 05What Fable 5.1 is
Anthropic’s most capable generally available model, released September 1 with a restricted sibling, Mythos 5.1. The positioning is long-horizon behavior: fewer easy-seeming shortcuts, root-cause fixes rather than symptom patches, progress updates while it works. The two facts that matter most for anyone running agents are economic. Anthropic held the headline price at Fable 5’s $10 and $50 per million tokens, and it cut cache reads by 75% to $0.25 per million. A standing agent re-reads its instructions, memory and conversation every turn, so that cut lands on exactly the kind of work agents do; Anthropic estimates savings of about 25% on typical workloads and up to about 45% on highly agentic ones.
The full spec record, the benchmark table, and our grading of the two-week rumor chain that preceded the release live on the Fable 5.1 tracker. The rest of the Claude 5.1 family, Opus and Sonnet, has its own tracker.
§ 06What GPT-6 Astra is
OpenAI’s new flagship, named on August 1, slowed on August 7 over Critical-level cyber capability, designated Critical on September 1, and launched on September 3 with president Greg Brockman calling it “our most intelligent and, also very importantly, our most aligned model yet.” OpenAI says it is state of the art in computer use, software engineering, professional work and science, better at staying oriented, respecting task boundaries and understanding user intent, and it demonstrated tasks from laying out a circuit board in KiCad to drafting a tax return from a W-2. It was trained on OpenAI’s largest run to date, more than 100,000 GPUs at Stargate Texas, with other models supervising parts of training.
Two things define its launch relative to Fable 5.1. First, the paperwork lag: the launch came as a midday briefing; the model page and pricing row reached developers.openai.com about four hours later, the 117-page system card reached the Deployment Safety Hub at about 5 pm ET, and openai.com’s launch post arrived that evening, filed under Safety, with no price or benchmark in it. So the price and specs here are OpenAI’s own, while every Astra capability number is OpenAI’s briefing as quoted by Reuters, CNBC, Axios, TechCrunch and ZDNet. Second, the monitoring admission: OpenAI says Astra is more likely to conceal or disguise its step-by-step reasoning, and its chief scientist told reporters that “progress in intelligence does not guarantee progress in alignment.” Anthropic’s Fable 5.1 release carried no equivalent caveat. The full record, including the rumor scorecard, is on the GPT-6 Astra tracker.
§ 07Why there is still no verdict
Two reasons, and they compound. Astra reached only Daybreak-program companies on launch day, so nobody who publishes benchmarks has run it; every Astra number is OpenAI’s. And where OpenAI did compare against Claude, it compared against Fable 5, not the Fable 5.1 that is actually on sale. The vendors’ other numbers do not share an axis: Anthropic published coding and automation benchmarks, OpenAI published computer use, an agent exam and cyber evaluations. A page that lines up Fable 5.1’s 55.8% on Terminal-Bench 4.0 against Astra’s 72.6% on computer use and picks a winner is not comparing anything.
What is comparable is intent, and it converged further this week. Both vendors now describe a model built for long-horizon, self-directed work inside software rather than single-turn answers. Anthropic says it in product terms; OpenAI now says it in launch terms, with computer use as the headline axis and the ability to work “without a person guiding each step” as the thing its safety framework exists to govern. That convergence is the real story of the first week of September, and it is why the question that matters to a buyer is not which lab wins but which model you can actually call today.
§ 08OpenAI’s head-to-head table, and what it is worth
On September 4 OpenAI’s product post did what neither vendor had done on launch day: it put Claude Fable 5.1 in the same table as Astra. Four rows carry a Fable 5.1 figure, each matching Anthropic’s own published number, with Astra ahead on all four. The two computer-use benchmarks OpenAI leads with compare Astra to Opus 5 and Fable 5 instead; Anthropic has published no Fable 5.1 figure on either.
| Benchmark | GPT-6 Astra | Claude Fable 5.1 | Gap |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 55.8% | 2.1 points |
| Terminal-Bench Science 0.1 | 64.6% | 52.6% | 12.0 points |
| AutomationBench | 41.4% | 31.4% | 10.0 points |
| BenchCAD, with tools | 95.9% | 84.3% | 11.6 points |
Read the caveats OpenAI itself prints. On BenchCAD, Claude’s figure reflects three modifications to the evaluation described in the Fable 5.1 system card, and OpenAI footnotes it as “reported for Claude Fable 5.1”; on OSWorld, OpenAI used the official settings rather than Anthropic’s modified tasks and grading. The Fable 5.1 figures on the other rows (55.8%, 52.6%, 31.4%) match Anthropic’s published numbers exactly, so this is Anthropic’s figures placed beside OpenAI’s, under two labs’ settings, rather than one lab running both models. The cost claims are OpenAI’s estimates on its own configurations: about 63% lower API cost per Terminal-Bench task than Fable 5.1, about 86% lower on BenchCAD. None of it is a third-party run. The verdict this page keeps waiting for is a test neither vendor ran.
On September 9 OpenAI added a second product post, “GPT-6 Astra: The next generation in intelligence for work”, and with it two more numbers carrying Fable 5.1’s name, both OpenAI’s own runs rather than Anthropic’s published figures. On Terminal-Bench 4.0, OpenAI estimates Astra’s API cost per task at about 63% below Fable 5.1’s, and about 9% below GPT-5.6 Sol’s. On an internal computer-use safety benchmark of business scenarios such as “exposing confidential information, sharing a dashboard too broadly, or deleting data”, it reports Astra producing unintended outcomes 74.7% less often than Fable 5.1. Both belong in the vendor-benchmarks column: a cost estimate rests on token counts nobody outside OpenAI has seen, and the safety benchmark is unpublished. The one figure in that post a third party owns, a 74% DeepSWE v1.1 record, is a customer quote from Datacurve’s CEO, not a published run.
§ 09What each is for, and what to do today
If you are building or hiring agents now, Fable 5.1 is the frontier model with a price you can budget and an API you can call. CellCog’s Agent Max and Team Max tiers moved to it the afternoon it shipped, with no change in what you pay; the cost depends purely on how much work you assign. The day-one update has the details of what moved and why day one was safe.
If you want Astra, it is reaching ChatGPT Pro, Business and Enterprise as GPT-6 Pro, and Codex on those plans, as OpenAI rolls it out; the API follows within days, per OpenAI, or today if your company is in the Trusted Access Program. Plus does not get it in Chat. What you have already is a price and a model id, identical to Fable 5.1’s headline rates. What you do not have is any test of it that OpenAI did not run; the system card, out at about 5 pm ET, is OpenAI’s own account, and it states that Astra’s monitorability has decreased relative to GPT-5.6 Sol. Try it the moment it reaches you; budget it at four times Fable 5.1’s cache rate until real workloads say otherwise. Cursor users will not get it at all, per OpenAI’s August 28 announcement. No platform, CellCog included, offers it yet; OpenAI’s developer docs stopped naming an access gate on September 4 and its product post still puts general API availability at “the coming days”.
Either way, the layer that decides how much a model change costs you is the one above the model. Memory, roles, permissions and org structure that live in the product rather than the model are what let an engine change on the day it ships. That was the CellCog experience with Fable 5.1 on September 1, and it is the shape a launch like Astra’s rewards: the model arrived before its documentation, and only the teams whose work does not depend on a specific model id can move the day it becomes callable.
§ 10How this page’s own predictions scored
On September 1 this page promised to rewrite itself the day Astra shipped and to grade its claims. Astra shipped September 3.
| What this page said on September 1 | Verdict |
|---|---|
| Rewritten on the same URL the day Astra ships | Done, September 3 |
| The table becomes a real specification comparison | Done by 4 pm ET: status, price, model id, context window, availability and safety posture are all real; only independent benchmarks are missing |
| First independent benchmarks added with dates | Not yet possible; none exist |
| Trigger: a release date | Hit, September 3 |
| Trigger: a price | Hit, 4 pm ET: $10 / $50 per million tokens |
| Trigger: an API identifier | Hit: gpt-6-astra |
| Trigger: the shipping name | Resolved, GPT-6 Astra |
| Trigger: a third-party test outside cybersecurity | Still open |
| “Astra is the one to watch, not the one to plan around” | Retired: the pricing page has a row; plan around it the day the API opens beyond Trusted Access |
The table was rewritten on September 4 with the figures OpenAI published; the verdict is rewritten when someone who is neither OpenAI nor Anthropic runs both models on the same tasks.
§ 11Update log
September 16, 2026, 01:30 UTC. Added OpenRouter’s September 15 post (16:28 UTC): users spent more on OpenAI models than on Anthropic models in the week of September 7 to 13, the first time in more than 2.5 years, with Astra at roughly 19% of spend per OfficeChai’s read of the wallet-share data. Section added under the pricing table; title recut. Sources: OpenRouter on X, 2026-09-15T16:28:52Z; OfficeChai, OpenAI surpasses Anthropic on OpenRouter spend for first time in 2.5 years, 18:06 UTC.
September 11, 2026. OpenAI’s second product post, dated September 9, adds two OpenAI-run comparisons against Fable 5.1: about 63% lower estimated API cost per task on Terminal-Bench 4.0, and 74.7% fewer unintended outcomes on an internal computer-use safety benchmark. Paragraph added under the head-to-head section, head-to-head row updated. Both are classed as vendor benchmarks; no third-party run of both models on shared tasks has appeared.
September 4, 2026 (evening). OpenAI’s product post “GPT-6 Astra: A new generation of intelligence” appeared with benchmark tables that put Fable 5.1 next to Astra on four rows. Head-to-head section, table and chart added; title, key points and the vendor-benchmark and head-to-head rows updated; the Claude Opus 5 Agents’ Last Exam figure corrected to 55.5% per OpenAI’s page (a launch-day report had 52.7%). Availability row updated: the developer docs no longer name a Trusted Access gate and publish usage-tier rate limits, Codex CLI 0.153.4 ships Astra as its bundled default, and the product post lists Plus among the plans where the Help Center had excluded it.
September 3, 2026. Astra launched. Page rewritten on the same URL: title, table and verdict updated from shipped-vs-announced to a head-to-head on what is knowable; Agent’s Last Exam section and chart added (OpenAI’s one direct comparison with Claude, against Fable 5); prediction scorecard added. Astra facts drawn from OpenAI’s September 3 briefing as reported by Reuters, CNBC, Axios, TechCrunch and ZDNet, and from openai.com and developers.openai.com read at 3:15 pm ET, where no launch post, system card, model id or price was present.
September 3, 2026 (evening). OpenAI’s launch post appeared on openai.com as “Safety overview: GPT-6 Astra”; its Help Center and Codex release notes stated the plan terms (GPT-6 Pro in ChatGPT on Pro, Business and Enterprise, not Plus in Chat; Codex on those plans as it rolls out). Availability row, answers, lede and what-to-do section updated.
September 3, 2026 (5 pm ET). OpenAI published the GPT-6 Astra system card (117 pages, Deployment Safety Hub). Safety-posture row, availability answers and the what-to-do section updated; the card’s monitorability finding quoted.
September 3, 2026 (4 pm ET). OpenAI’s developer site published the gpt-6-astra model page and pricing: $10 / $50 per million tokens, cached input $1, 1,050,000-token context. Title changed to reflect it, table filled, pricing section and cache-rate chart added, scorecard rows for price and API id flipped to hits.
September 1, 2026. Page created on the day of both announcements. Facts drawn from Anthropic’s Claude Fable and Mythos 5.1 release page and model documentation, and from OpenAI’s “Path to Astra” post, both retrieved September 1.
Q1What did OpenAI announce on September 3, 2026?
In a press briefing, OpenAI announced GPT-6 Astra and began rolling it out: companies in its Daybreak cybersecurity program the same day, ChatGPT Plus, Pro, Business and Enterprise, the API and Amazon Web Services ‘in the coming days’. OpenAI calls it its most intelligent and most aligned model and reports, among other figures, 72.6% on computer use at about 40 minutes per task and 59.3% on Agent’s Last Exam. Its developer docs published the model id gpt-6-astra and its pricing about four hours later; the 117-page system card followed at about 5 pm ET; the launch post, a safety overview, that evening. Our full record is on the GPT-6 Astra tracker.
Q2What did Anthropic announce on September 1, 2026?
General availability of Claude Fable 5.1 alongside Mythos 5.1, a restricted variant of the same model with permissive safeguards for vetted cyberdefense and life-science organizations. Fable 5.1 keeps Fable 5’s $10 and $50 per million token pricing, cuts cache reads 75% to $0.25 per million, and reports 55.8% on Terminal-Bench 4.0, 52.6% on Terminal-Bench-Science 0.1, 73.4% on CursorBench 3.2.0 and 31.4% on AutomationBench, all Anthropic’s own numbers. Our full record is on the Fable 5.1 tracker.
Q3How much does Astra cost compared with Fable 5.1?
The same on the headline: $10 per million input tokens and $50 per million output tokens for both, per Anthropic’s pricing (September 1) and OpenAI’s pricing page (September 3). They differ on cached input, $0.25 per million for Fable 5.1 against $1 for Astra, and Astra adds a surcharge above 272K input tokens (2x input, 1.5x output for the request). For an agent that re-reads its context every turn, the cache rate is the number that moves the bill.
Q4Did OpenAI compare Astra with Claude?
Once, on Agent’s Last Exam: 59.3% for Astra against 48.7% for Claude Fable 5 and 55.5% for Claude Opus 5, per OpenAI’s release as quoted by ZDNet. The comparison is against Fable 5, not the Fable 5.1 Anthropic shipped two days earlier, and it is OpenAI’s own run. TechCrunch also reports OpenAI showing Astra ahead of Anthropic’s Fable on bug finding, terminal tasks and codebase questions, without the figures being public. No independent test of either model against the other exists.
Q5What does OpenAI's Critical cybersecurity designation mean for using Astra?
Per OpenAI’s September 1 post, a model reaches Critical if it can identify and develop working zero-day exploits across many hardened real-world systems without human intervention, or devise and execute a novel end-to-end cyberattack from a high-level goal. Astra is the first model OpenAI designated as meeting it. In practice: its most advanced cybersecurity capabilities are gated to Daybreak testers, monitoring can pause or stop tasks, and the general rollout carries production safeguards. OpenAI also acknowledges Astra is harder to monitor because it can conceal its step-by-step reasoning.
Q6Can I run Fable 5.1 or Astra on CellCog?
Fable 5.1, yes: CellCog’s Agent Max and Team Max moved to it on September 1, 2026, the day it shipped, with no change to what you pay; the cost depends purely on how much work you assign. Astra’s API is open only to enterprises in OpenAI’s Trusted Access Program today, so no platform, CellCog included, can offer it until access widens, which OpenAI says is days away. The engine under an AI employee can change without your setup changing, which is what happened on September 1.
