Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact

Muse Spark 1.3 Is Out: Meta Claims Frontier Parity, and Its Own Table Mostly Backs It

At a glanceQuick answers
What is Muse Spark 1.3?
Meta’s newest proprietary frontier model, released September 2, 2026, trained for long-horizon agentic work and coding, with text, image, and video input and a 1M-token context window.
What does it cost?
$1.25 per million input tokens and $4.25 per million output tokens ($0.15 cached input) on the Meta Model API. A contributor tier at $0.10 and $0.20 lets Meta use your traffic to improve its products.
How does it compare to Claude Opus 5 and GPT-5.6 Sol?
On Meta’s own table, level with Opus 5 on agentic and professional work, ahead of both on long-horizon coding and long-context retrieval, behind both on agentic browsing and instruction following. Independent index: sixth overall.
Can I use the max mode?
Not yet. Meta says max reasoning arrives after additional safety testing. Muse Code and the Model API serve the previously available reasoning modes today.
Hand-drawn diagram of a three-step version ladder labeled 1.1, 1.2, and Muse Spark 1.3 with a dial reading MAX, two doorways labeled $1.25 / $4.25 and $0.10 / $0.20 contributor, and a terminal window labeled Muse Code pointing to a cloud labeled Model API
Fig 0Fourth Muse Spark release in five months, two doors in: standard, or the cheaper door where your prompts train Meta's models.

Meta released Muse Spark 1.3 on September 2, 2026, and for the first time its own table puts a Muse model level with Claude Opus 5 on agentic work. It is the fourth Muse Spark release since the line debuted in April, shipping the same day in Muse Code and the Meta Model API. The headline claims come with two caveats worth knowing before you read a single number: the benchmark column is the max reasoning mode, which Meta says is not public yet, and the numbers are Meta’s runs, under a methodology Meta published alongside them.

This page is built the way our Gemini 3.8 Flash record and our Fable 5.1 record were: from the vendor’s own pages, read the day they went up, with vendor-run numbers labeled as such and the independent read placed next to them.

On this page · 9 sectionsOpen
  1. What shipped
  2. The benchmarks: Meta’s full table
  3. Two vendor tables, one week apart
  4. The pricing, and the second door
  5. What Meta says about behavior and safety
  6. The max mode that is not shipping yet
  7. Who should care, and what to do
  8. What it means for CellCog users
  9. The record
Key points7 · 14 min full read
  1. Meta released Muse Spark 1.3 on September 2, 2026, its fourth Muse Spark release since April, available the same day in Muse Code and through the Meta Model API. Meta says it will also roll out to the Meta AI app, Facebook, and Instagram in the coming days.
  2. Meta’s published table compares Muse Spark 1.3 (max) with Muse Spark 1.2, GPT-5.6 Sol, and Claude Opus 5 across 11 benchmarks. It leads or ties on DeepSWE v1.1 (75.4%), SWE-Atlas (59.4%), Terminal-Bench 2.1 (88.8%), and both long-context MRCR bands (98.5% and 98.1%), and sits within a point of Opus 5 on JobBench, OSWorld 2.0, and AutomationBench.
  3. The column with those numbers is max reasoning, and Meta’s announcement says max is ‘coming shortly after we finish additional safety testing.’ Previously available reasoning modes ship today; the headline mode does not yet.
  4. Pricing on Meta’s model page: $1.25 per million input tokens, $0.15 cached, $4.25 output, with a 1M-token context window. A second tier, muse-spark-1.3-contributor, costs $0.10 input and $0.20 output, and Meta labels it ‘used to improve our products.’ Meta’s chief AI officer said developers pay no more than they did for 1.2.
  5. Independent read: Artificial Analysis scores Muse Spark 1.3 (max) at 62 on its Intelligence Index, sixth of 636 models, behind Claude Fable 5.1 and Claude Opus 5 and ahead of GPT-5.6 Sol.
  6. Meta reports about 20% fewer tool calls and about 25% fewer tokens than 1.2 on the same coding tasks, and says the model now confirms before consequential actions and asks clarifying questions when a prompt is ambiguous.
  7. Read the vendor tables side by side: Meta’s Opus 5 numbers differ from the ones Google published for the same model 24 hours earlier, because vendors pick harnesses, versions, and effort levels. Same-vendor generation deltas are the reliable column.

§ 01What shipped

Meta’s announcement leads with agentic workflows: a model trained across “a diverse set of harnesses” to sustain long tasks, juggle several workflows in one thread, ask clarifying questions when a prompt is ambiguous, and confirm before consequential actions. The model page carries the pricing table and the benchmark graphic.

Item Detail
Announced September 2, 2026, Meta AI Research blog
Available in Muse Code (macOS, Linux) and Meta Model API, same day
Consumer rollout Meta AI app, Facebook, Instagram, “in the coming days” (Wang, via Bloomberg)
Reasoning modes Previously available modes today; max reasoning after additional safety testing
Context window 1M tokens
Inputs and output Text, image, video in; text out
Standard price $1.25 input / $0.15 cached input / $4.25 output per 1M tokens
Contributor price $0.10 input / $0.002 cached / $0.20 output per 1M tokens, “used to improve our products”
Efficiency vs 1.2 About 20% fewer tool calls, about 25% fewer tokens on comparable coding tasks (Meta engineers’ comparisons)
License Proprietary; parameter count undisclosed
Roadmap, in Meta’s words “Bigger models, the Muse Spark open weights release, and more”
Table 1Muse Spark 1.3: the shipped record (September 2, 2026)

The release cadence is the first thing to notice. Muse Spark launched in April, 1.1 arrived in July with the first paid API access, 1.2 shipped August 5 alongside Muse Code, and 1.3 landed four weeks later. That is a monthly cadence from the lab that, a year ago, was releasing open Llama weights on a much slower clock.

§ 02The benchmarks: Meta’s full table

Meta publishes an 11-row comparison against three models: its own Muse Spark 1.2, OpenAI’s GPT-5.6 Sol, and Anthropic’s Claude Opus 5. The evaluation methodology states the rules: max effort for Muse Spark 1.3, Opus 5, and Sol; xhigh for 1.2; for each model, the highest comparable primary-metric value from Meta’s evaluation, the official leaderboard, or the provider’s self-reported result. Refusals and ungradable answers score zero and stay in the denominator.

Benchmark Muse Spark 1.3 (max) Muse Spark 1.2 (xhigh) GPT-5.6 Sol (max) Claude Opus 5 (max)
GDPVal-AA v2 (knowledge work, Elo) 1754 1615 1710 1824
JobBench (professional tool use) 64.9 61.6 45.4 65.7
OSWorld 2.0 (agentic computer use) 66.9 47.6 62.7 68.3
DeepSearchQA (agentic browsing, F1) 89.4 85.9 93.0 90.4
Agentic IF Index (Meta internal) 57.8 46.2 60.5 59.1
AutomationBench (business workflows, pass@1) 49.4 38.2 46.7 50.3
MRCR v2, 256K to 512K (long-context retrieval) 98.5 66.3 91.5 not reported
MRCR v2, 512K to 1M 98.1 55.5 73.8 not reported
DeepSWE v1.1 (long-horizon agentic coding) 75.4 55.0 73.0 74.0
SWE-Atlas Codebase QnA 59.4 46.2 53.5 52.7
Terminal-Bench 2.1 (agentic terminal coding) 88.8 82.9 88.8 86.7
Table 2Meta’s published comparison: Muse Spark 1.3 (max) against three models (September 2, 2026)

Three readings, one chart each.

First, the generation jump. Every row moved up from 1.2, and the moves are large where 1.2 was weakest: retrieval deep into the context window, long-horizon coding, and driving a desktop.

Muse Spark 1.3 vs 1.2: eight rows from Meta's tableDot chart of eight benchmarks with Muse Spark 1.3 highlighted ahead on every row: MRCR 512K to 1M 98.1 vs 55.5, MRCR 256K to 512K 98.5 vs 66.3, Terminal-Bench 88.8 vs 82.9, DeepSearchQA 89.4 vs 85.9, DeepSWE 75.4 vs 55.0, OSWorld 66.9 vs 47.6, JobBench 64.9 vs 61.6, AutomationBench 49.4 vs 38.2Muse Spark 1.3Muse Spark 1.2MRCR 256K to 512KMRCR 512K to 1MDeepSearchQATerminal-Bench 2.1DeepSWE v1.1OSWorld 2.0JobBenchAutomationBench098.5Muse Spark 1.3 vs 1.2: eight rows from Meta's tableDot chart of eight benchmarks with Muse Spark 1.3 highlighted ahead on every row: MRCR 512K to 1M 98.1 vs 55.5, MRCR 256K to 512K 98.5 vs 66.3, Terminal-Bench 88.8 vs 82.9, DeepSearchQA 89.4 vs 85.9, DeepSWE 75.4 vs 55.0, OSWorld 66.9 vs 47.6, JobBench 64.9 vs 61.6, AutomationBench 49.4 vs 38.2Muse Spark 1.3Muse Spark 1.2MRCR 256K to 512KMRCR 512K to 1MDeepSearchQATerminal-Bench 2.1DeepSWE v1.1OSWorld 2.0JobBenchAutomationBench098.5
Fig 1Muse Spark 1.3 vs 1.2: eight rows from Meta's table

Second, the parity claim. Put 1.3 next to Opus 5 and Sol on the seven rows where all three have a score, and the claim Alexandr Wang made to Bloomberg, “competitive” with Anthropic’s best and “better than” OpenAI’s, reads as mostly true on Meta’s own numbers. Against Opus 5 it is within a point on JobBench, AutomationBench, and DeepSearchQA, within 1.4 on OSWorld, and ahead on the three coding rows. Against Sol it wins everything except agentic browsing and Meta’s own instruction-following index. The one clear gap is GDPVal-AA, where Opus 5 holds a 70-point Elo lead.

Muse Spark 1.3 against Claude Opus 5 and GPT-5.6 Sol on seven rowsDot chart with Muse Spark 1.3 highlighted: ahead on DeepSWE (75.4, 74.0, 73.0), SWE-Atlas (59.4, 52.7, 53.5) and Terminal-Bench (88.8, 86.7, 88.8), within a point of Opus 5 on JobBench (64.9, 65.7, 45.4), OSWorld (66.9, 68.3, 62.7) and AutomationBench (49.4, 50.3, 46.7), behind both on DeepSearchQA (89.4, 90.4, 93.0)Muse Spark 1.3Claude Opus 5GPT-5.6 SolDeepSearchQATerminal-Bench 2.1DeepSWE v1.1OSWorld 2.0JobBenchSWE-Atlas Codebase QnAAutomationBench093Muse Spark 1.3 against Claude Opus 5 and GPT-5.6 Sol on seven rowsDot chart with Muse Spark 1.3 highlighted: ahead on DeepSWE (75.4, 74.0, 73.0), SWE-Atlas (59.4, 52.7, 53.5) and Terminal-Bench (88.8, 86.7, 88.8), within a point of Opus 5 on JobBench (64.9, 65.7, 45.4), OSWorld (66.9, 68.3, 62.7) and AutomationBench (49.4, 50.3, 46.7), behind both on DeepSearchQA (89.4, 90.4, 93.0)Muse Spark 1.3Claude Opus 5GPT-5.6 SolDeepSearchQATerminal-Bench 2.1DeepSWE v1.1OSWorld 2.0JobBenchSWE-Atlas Codebase QnAAutomationBench093
Fig 2Muse Spark 1.3 against Claude Opus 5 and GPT-5.6 Sol on seven rows

Third, the row Meta wants you to see. The long-context numbers are the outliers in the whole table. On MRCR v2’s hardest band, 512K to 1M tokens with 8 needles, Muse Spark 1.3 scores 98.1 against 73.8 for GPT-5.6 Sol, and Meta reports no Opus 5 figure at all. If it holds up in independent runs, that is the difference between a model that can use a million-token window and a model that merely accepts one.

MRCR v2, 512K to 1M tokens: retrieval deep in the context windowBar chart with Muse Spark 1.3 highlighted at 98.1, GPT-5.6 Sol at 73.8, Muse Spark 1.2 at 55.5Muse Spark 1.398.1GPT-5.6 Sol73.8Muse Spark 1.255.5MRCR v2, 512K to 1M tokens: retrieval deep in the context windowBar chart with Muse Spark 1.3 highlighted at 98.1, GPT-5.6 Sol at 73.8, Muse Spark 1.2 at 55.5Muse Spark 1.398.1GPT-5.6 Sol73.8Muse Spark 1.255.5
Fig 3MRCR v2, 512K to 1M tokens: retrieval deep in the context window

§ 03Two vendor tables, one week apart

Here is a check anyone can run. Google published its own comparison table on September 2 for Gemini 3.8 Flash, and it also scored Claude Opus 5. On three benchmarks both vendors ran, the Opus 5 numbers do not match.

Benchmark Opus 5 per Google (Sep 2) Opus 5 per Meta (Sep 2) Muse Spark 1.3 per Meta Gemini 3.8 Flash per Google
DeepSWE v1.1 74.0 74.0 75.4 73.7
OSWorld 2.0 75.4 68.3 66.9 59.0
Terminal-Bench 2.1 89.1 86.7 88.8 89.4
Table 3The same third-party model, scored by two vendors in one week

DeepSWE agrees because both vendors took it from Datacurve’s public leaderboard. OSWorld disagrees by seven points; Meta’s methodology notes it ran version 08.08 in an internal framework with its own computer-control tool, and Google’s page reports a partial score under its own setup. Terminal-Bench disagrees by 2.4 points, and Meta’s note says it ran each task “with the coding agent named in the result,” while Sol’s score came straight from OpenAI’s model card. None of this is bad faith. It is what happens when every vendor picks harnesses, versions, and effort levels, then reports the highest comparable number it can find. The practical rule: trust a vendor’s generation-over-generation deltas, treat its cross-vendor rows as a range, and wait for the independent index.

That index exists. Artificial Analysis scores Muse Spark 1.3 (max) at 62 on its Intelligence Index, sixth of 636 models, behind Claude Fable 5.1 (high) at 62.5 and Claude Opus 5 at 63.1, and ahead of GPT-5.6 Sol (max) at 60.9. The xhigh variant scores 60.8. That is a narrower and more defensible version of Meta’s claim: not ahead of the frontier, but on it, at a quarter to a fifth of Opus 5’s output price.

§ 04The pricing, and the second door

The standard tier is straightforward, and Wang told Bloomberg developers pay no more for 1.3 than they did for 1.2. Against the models in Meta’s own table, the output price is the argument.

Model Input Output
Muse Spark 1.3 (contributor) $0.10 $0.20
Muse Spark 1.3 (standard) $1.25 $4.25
GPT-5.6 Sol $4.00 $20.00
Claude Opus 5 $5.00 $25.00
Table 4Published output price per 1M tokens, the models in Meta’s comparison
Output price per 1M tokens, the models in Meta's tableBar chart of output prices with Muse Spark 1.3 standard highlighted at 4.25, contributor at 0.20, GPT-5.6 Sol at 20, Claude Opus 5 at 25Muse Spark 1.3 (contributor)0.20Muse Spark 1.3 (standard)4.25GPT-5.6 Sol20Claude Opus 525Output price per 1M tokens, the models in Meta's tableBar chart of output prices with Muse Spark 1.3 standard highlighted at 4.25, contributor at 0.20, GPT-5.6 Sol at 20, Claude Opus 5 at 25Muse Spark 1.3 (contributor)0.20Muse Spark 1.3 (standard)4.25GPT-5.6 Sol20Claude Opus 525
Fig 4Output price per 1M tokens, the models in Meta's table

The second door deserves a plain reading. Meta’s pricing table lists two model ids. muse-spark-1.3 at $1.25 and $4.25 carries the label “Not used to improve our products.” muse-spark-1.3-contributor at $0.10 and $0.20 carries the label “Used to improve our products.” That is roughly a 92% discount in exchange for your prompts and outputs becoming training material. Meta is not hiding it; the two labels sit one row apart. For a hobby project it is close to free. For anything touching customer data, contracts, or code you would not publish, the standard tier is the only one to consider, and the discount is the price of that rule being easy to forget at 2am.

The efficiency claim matters as much as the unit price, and it points the opposite way from Google’s caveat a day earlier. Where Google warned that Gemini 3.8 Flash “works harder” and may spend more tokens per task, Meta says 1.3 spends fewer: about 20% fewer tool calls and 25% fewer tokens than 1.2 in comparisons by its own engineers. If that holds on your workload, the same job costs less on 1.3 than on 1.2 at an unchanged price. As with Google’s claim, the right move is a measurement on your real traffic, not a migration on a press release.

§ 05What Meta says about behavior and safety

Three claims in the announcement bear directly on anyone handing this model real work.

  • Confirmation before consequential actions. Meta says 1.3 has “better calibration on what constitutes irreversible actions and proceeds accordingly,” and Wang put it more simply: when it is about to do something irreversible, it asks. That is the right default for an agent with tool access, and it is a product decision as much as a model property.
  • Clarifying questions and asking for help. The model “asks clarifying questions when prompts are ambiguous” and “invokes help from the user when stuck,” rather than guessing. Meta also says it has a better sense of what it does and does not know, and flags hurdles instead of hallucinating outcomes.
  • Prompt-injection resistance, without a number. Meta claims “stronger adversarial robustness, with improved resistance to adversarial inputs and prompt injections” but publishes no figure. For scale, Google’s Gray Swan chart from September 2 put Muse Spark 1.2 at a 24.2% attack success rate within 15 attempts, against 4.8% for Claude Opus 5 and 5.5% for Gemini 3.8 Flash. Whether 1.3 closed that gap is exactly the number this release does not give you.

§ 06The max mode that is not shipping yet

The announcement’s availability paragraph is careful: “Previously available reasoning modes are available today with max reasoning coming shortly after we finish additional safety testing.” Every number in the comparison table is from the max mode. Artificial Analysis describes Muse Spark 1.3 (max) as being in limited preview for Meta’s partners, and scores the publicly available xhigh mode more than a point lower on its index.

So the model you can call today through Muse Code or the API is not the model in the table. That is a normal staging pattern, and Meta says the gap is safety testing, not capability. It still means a developer benchmarking 1.3 this week should expect xhigh-class results, and the max column becomes real on the day Meta flips it on. We will update this page that day.

§ 07Who should care, and what to do

If you build on Muse Spark 1.2 or Muse Code: 1.3 is already the default in Muse Code, and the API id changes to muse-spark-1.3. Re-run your evals at the effort level you actually use, check the token counts Meta says should fall, and read the two pricing labels before you pick a model id.

If you are choosing an AI agent or AI employee platform: count the week. Fable 5.1 on September 1, Gemini 3.8 Flash and Qwen3.8-Max-0902 on September 2, Muse Spark 1.3 the same day, GLM-5.3-Flash’s promotional pricing ending September 9. The model layer now rotates weekly. The question to ask a platform is not which model it runs today but whether the work you build on it survives the next swap: the roles, the memory, the accumulated context, the tasks in flight.

§ 08What it means for CellCog users

CellCog does not run on Muse Spark 1.3. Its Flash tiers route to Gemini 3.8 Flash, its Max tiers to Claude Fable 5.1, and its Creative tier to Claude Opus 5, each swapped in on the day the model shipped. Every frontier release, this one included, is evaluated against that routing, and if 1.3’s public modes earn a place, the switch happens underneath your employees without a change on your side. Your pricing is not tied to any of it: a full shift of real work runs about $25, and the cost depends purely on how much work you assign. How the tiers fit together is in Flash, Core, Max.

§ 09The record

As of September 3, 2026: Muse Spark 1.3 is available in Muse Code and the Meta Model API at $1.25/$4.25 per million tokens (contributor tier $0.10/$0.20); max reasoning is pending additional safety testing; Meta’s 11-row comparison table is published with its methodology; Artificial Analysis ranks the max preview sixth on its Intelligence Index; consumer rollout to Meta AI, Facebook, and Instagram is promised within days; open weights for Muse Spark are on Meta’s stated roadmap without a date. If max reasoning goes public, if independent replications land materially different from Meta’s numbers, if the consumer rollout completes, or if the open-weights release is dated, the update happens here, same day.

Frequently asked6 questions

Q1When was Muse Spark 1.3 released?

September 2, 2026, announced on Meta AI Research’s blog. It shipped the same day in Muse Code and the Meta Model API. Meta’s chief AI officer, Alexandr Wang, told Bloomberg it would reach the Meta AI app, Facebook, and Instagram in the coming days.

Q2What are the confirmed specifications?

From Meta’s model page and evaluation methodology: a 1M-token context window, text, image, and video input with text output, reasoning with multiple effort levels including a max level, and a proprietary license. Meta has not disclosed parameter count. Artificial Analysis lists it as a reasoning model released September 2, 2026.

Q3How does Muse Spark 1.3 compare to Muse Spark 1.2?

Every row on Meta’s table moved up. The largest jumps are long-context retrieval (MRCR 512K to 1M: 98.1% vs 55.5%), agentic coding (DeepSWE v1.1: 75.4% vs 55.0%), and computer use (OSWorld 2.0: 66.9% vs 47.6%). Meta also reports about 20% fewer tool calls and 25% fewer tokens on comparable coding tasks, so a task should cost less even at an unchanged unit price.

Q4What is the contributor tier?

A second API price point, muse-spark-1.3-contributor, listed at $0.10 per million input tokens, $0.002 cached, and $0.20 output, roughly one twelfth of the standard tier. Meta’s own label on the pricing table is ‘used to improve our products,’ against ‘not used to improve our products’ on the standard tier. It is a data-for-discount trade, stated plainly.

Q5Are the benchmark numbers independent?

No. They are Meta’s runs, under a methodology Meta published: max effort for Muse Spark 1.3, Opus 5, and GPT-5.6 Sol, xhigh for Muse Spark 1.2, and for each model the highest comparable score from Meta’s own evaluation, the official leaderboard, or the provider’s self-reported result. Refusals score zero. The one internal benchmark (Agentic IF Index) has no public task set. Artificial Analysis’s Intelligence Index is the independent cross-check.

Q6Does CellCog run on Muse Spark 1.3?

No. CellCog’s Flash tiers run on Gemini 3.8 Flash, its Max tiers on Claude Fable 5.1, and its Creative tier on Claude Opus 5. Every frontier release is evaluated against that routing, and the work you build on CellCog, the roles, memory, and tasks, does not depend on which model sits underneath. A full shift of real work runs about $25, and the cost depends purely on how much work you assign.

Published 03 September 2026 All Choosing a platform →