On September 2, 2026, Meta released "Muse Spark 1.3," an update strengthening long-running agentic work and coding, through Muse Code and the Meta Model API. The update comes just 28 days after the previous generation, 1.2, was announced. In benchmarks for long-context retrieval and code understanding, the model posts figures that rival or exceed GPT-5.6 Sol and Opus 5, and the standard API tier—which isn't used for product improvement—is priced lower than both competing models. However, the "max" reasoning tier that Meta placed as the primary column in its comparison table is still undergoing additional safety testing; only the previously existing reasoning tier is available at launch. Taking the published figures at face value as an "overall lead among current models" risks misjudging real-world performance at deployment.

AD

What changed in long-running task behavior over 28 days

Version 1.2, co-trained with Muse Code, debuted on August 5 as a model built to sustain long stretches of work from planning through implementation and verification. Version 1.3 extends that training focus further into long-duration coding work. According to Meta, the model now assembles necessary context from materials containing contradictions, corrects gaps in its plan mid-task, and correctly maps requests to the right task even when multiple jobs are interleaved within a single long conversation.

Interaction with users was also part of the training. The model now asks questions when instructions are ambiguous, requests help when it can no longer proceed on its own, and confirms before taking irreversible actions. This is a change aimed at reducing the risk—inherent in long-running agents—of compounding errors built on a mistaken premise. That said, the release materials do not provide numeric figures for the appropriateness of the model's questions or the rate at which it overlooks critical actions.

According to Meta engineers' comparison with 1.2, version 1.3 reduced tool calls by roughly 20% and token consumption by roughly 25%. Since the API bills by token volume, reproducing these gains would also affect the cost of completing a single task. However, the task set, number of trials, absolute values, and variance behind these figures are not disclosed. This is an internal Meta comparison against the prior version, not a claim that the model is 20–25% more efficient than GPT-5.6 Sol or Opus 5.

The xhigh tier available at launch is not the same as max-tier performance

Meta's evaluation table separates the not-yet-available max tier for 1.3 from the currently accessible xhigh tier. The comparison points are 1.2 xhigh, GPT-5.6 Sol max, and Opus 5 max. Figures are raw metrics for each evaluation, where higher is better. A dash ("―") indicates the value is not published.

Evaluation 1.3 max 1.3 xhigh 1.2 xhigh GPT-5.6 Sol max Opus 5 max
GDPVal-AA v2 1,754 1,709 1,615 1,710 1,824
JobBench 64.9 61.2 61.6 45.4 65.7
OSWorld 2.0 66.9 57.2 47.6 62.7 68.3
DeepSearchQA 89.4 89.4 85.9 93.0 90.4
Agentic IF Index 57.8 55.7 46.2 60.5 59.1
AutomationBench 49.4 43.8 38.2 46.7 50.3
MRCR 256K–512K 98.5 97.6 66.3 91.5
MRCR 512K–1M 98.1 93.1 55.5 73.8
DeepSWE v1.1 75.4 55.0 73.0 74.0
SWEAtlas CodeBase QnA 59.4 54.0 46.2 53.5 52.7
Terminal-Bench 2.1 88.8 89.2 82.9 88.8 86.7

Compared with max, the Muse Spark 1.3 xhigh tier available at announcement scores 9.7 points lower on OSWorld 2.0, 5.6 points lower on AutomationBench, 5.4 points lower on SWEAtlas CodeBase QnA, and 5.0 points lower on MRCR 512K–1M, while scoring 0.4 points higher on Terminal-Bench 2.1. Even between xhigh tiers, JobBench slipped from the previous 1.2's 61.6 to 61.2, showing the improvements are not uniform across categories.

Even on the current xhigh tier, the model surpasses GPT-5.6 Sol max on both MRCR context bands. On SWEAtlas, its score of 54.0 slightly exceeds GPT's 53.5 and Opus 5's 52.7, and on Terminal-Bench, its 89.2 also beats both companies. Conversely, Opus 5 leads on knowledge-work tasks in GDPVal-AA v2 and on computer-use tasks. GPT or Opus 5 top the rankings for search, instruction-following, and business automation. Long context and code-adjacent tasks are clear strengths for Muse Spark 1.3, but it does not lead across agentic work as a whole.

AD

Why a single table can't yield an overall ranking

The 11 rows do not share a common yardstick. GDPVal-AA v2 scores 220 deliverables across 44 occupations and nine U.S. industries using anonymous head-to-head comparisons, expressed as an Elo rating with humans fixed at 1,000. JobBench scores 65 tasks across 35 occupations against individual rubrics. AutomationBench judges 600 simulated business tasks by end-state outcome—adding or averaging 64.9 and 49.4 would be meaningless.

Execution conditions also differ row by row. DeepSearchQA measures 900 questions using the same search infrastructure and browser environment. MRCR embeds eight retrieval targets in 100 examples per context band and measures string-reproduction accuracy without tool use—these two evaluations are relatively easy to keep aligned across models. By contrast, DeepSWE v1.1 runs 1.3 through mini-swe-agent while pulling other models' figures from Datacurve's official leaderboard. Terminal-Bench 2.1 likewise uses a named coding agent or fixed environment specific to each model, with GPT's figures sourced from OpenAI's model card. These are not tests where the bare models were placed in an identical harness.

Meta further explains that for each model it selected the "best comparable figure" from its own runs, official leaderboards, or the developer's self-reported results. It also explicitly notes that the instructions, tools, and execution time used for third-party models may not reproduce each company's optimal configuration. OSWorld 2.0 uses version 06.24 for 1.2 only, while the others use 08.08—so the jump from 47.6 to 66.9 cannot be attributed purely to generational improvement. The Agentic IF Index is an internal Meta composite metric, and there is no single external task set for reproducing the same evaluation independently.

The published table is useful as a map for locating areas of strength. It cannot be used as an overall ranking table.

Pricing is a strong point, but the contributor tier has different data-use terms

On API pricing, Muse Spark 1.3 is clearly cheaper. Breaking down the public pricing for three models with roughly million-token context windows into new input, cached input, and output yields the following. Units are US dollars per million tokens.

Model / Tier Context Window Input Cached Input Output Terms on Public Page
Muse Spark Standard (1.3) 1M $1.25 $0.15 $4.25 Not used for product improvement
muse-spark-1.3-contributor 1M $0.10 $0.002 $0.20 Used for product improvement
GPT-5.6 Sol 1.05M $4 $0.40 $20 Limited-time pricing
Opus 5 1M $5 $0.50 $25 Standard API pricing

The standard Muse Spark 1.3 tier, which is not used for product improvement, prices at $1.25 per million tokens for input, $0.15 for cached input, and $4.25 for output—lower than GPT-5.6 Sol's $4, $0.40, and $20, and Opus 5's $5, $0.50, and $25. The contributor tier is priced at $0.10, $0.002, and $0.20, but Meta designates this tier as one where data is "used for product improvement."

If Meta's roughly 25% reduction in tokens versus the prior version is also reproduced, the cost gap for completing a given task could widen further. That said, actual billed amounts vary with retries, cache hit rates, tool fees, and each company's tokenizer. GPT-5.6 Sol's pricing is a limited-time offer valid at least through November 21, and for requests with input exceeding 272,000 tokens, the input price doubles and the output price rises 1.5x for the entire request.

The contributor tier is dramatically cheaper, but it is not simply a discount version under the same terms as the standard tier. Meta's table states only that data is "used for product improvement" without explaining how much of the code, prompts, and tool-call history is retained, for how long, or under what usage scope. Enterprises connecting confidential repositories need to verify the scope of covered data, retention period, and opt-out procedures in the contract documentation before considering price.

Gaps also remain in the availability terms on the Muse Code side. While the announcement states 1.3 is available through Muse Code, the current monthly plan table lists access to 1.2 under the $5 Everyday Usage tier and touts "access to the latest models" starting from the $15 High Usage tier. Whether 1.3 is selectable on the lowest-priced plan cannot be determined from this description alone.

AD

Adoption decisions should hinge on your own tasks, not the max tier's release date

For tasks involving extracting information from long repositories or understanding codebases, there is a reasonable case for including Muse Spark 1.3 as a candidate even at the currently available xhigh tier. Conversely, for search-centric research, GPT-5.6 Sol leads in Meta's table, and for knowledge-work deliverables or computer-use tasks, Opus 5 leads. Moreover, rows where the gap is only a few points come with no variance or confidence intervals reported, so they cannot be directly translated into operational superiority.

On safety, Meta states it has strengthened resistance to adversarial input and prompt injection, and improved judgment around irreversible actions. Since the release materials do not disclose the attack set, number of trials, or success rates, enterprises must verify this using their own tasks involving deletion, transmission, and permission changes. A future open-weight release is also only in preview stage, with model size, license terms, equivalence to 1.3, and release date all undetermined.

Once max becomes available, Meta's primary comparison column and the reasoning tier available in the actual product will align. However, the tasks and tools used in each evaluation will still differ, and the problem of differing execution environments and data sources persists row by row. When adopting the model, organizations should compare completion rates and total token counts against competitors in their own environment, measure the number of correction cycles required, confirm the rate at which the model correctly halts before critical actions, and determine whether they can achieve a per-task cost that matches the standard tier's data-use terms.