READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Meta's Muse Spark 1.3 Ties OpenAI and Grok on Benchmarks, But Its Best Score Comes From a Version Nobody Can Fully Access Yet

Meta's Muse Spark 1.3 Ties OpenAI and Grok on Benchmarks, But Its Best Score Comes From a Version Nobody Can Fully Access Yet
Meta rolled out Muse Spark 1.3 on Wednesday, September 2, with paid API access opening today, September 3. The publicly usable version ties OpenAI and Grok on independent rankings, but Meta's flashiest number belongs to a limited-preview variant with no public API provider. Zuckerberg's claims check out on the numbers he's not showing you.

Since Meta's first Muse Spark model launched April 8, 2026, the company has pushed out three follow-on versions in five months. Muse Spark 1.3 rolled into Meta's Muse Code terminal agent and the Meta Model API on Wednesday, September 2, with broader paid API access opening today, according to flowtivity.ai and BigGo Finance. It's the fourth release in the line and the one Meta is calling its biggest jump yet.

Mark Zuckerberg announced it on X, writing that Muse Spark 1.3 delivers "frontier performance almost too cheap to meter" and calling it Meta's "biggest jump we've made so far on coding and agentic work," according to VentureBeat and ExplainX.ai. That's a strong claim from the CEO. The numbers back some of it up, and complicate the rest.

Two Versions, One Headline Number

Meta shipped two configurations of 1.3. The one developers can actually use, called xhigh, scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol (max) and Grok 4.6 (high), according to Artificial Analysis, the independent benchmarking firm. That's up from 57 for Muse Spark 1.2 in August and 51 for 1.1 in July.

The other configuration, called max, scores 62, according to Artificial Analysis. But max is still in limited partner preview and completing safety testing, VentureBeat reported. Artificial Analysis says it evaluated max only in that limited preview and currently lists no API provider offering it at all.

Meta's launch materials prominently feature max's results, including a #1 score on Tau3-Bench Banking at 52% and a GDPval-AA v2 score of 1,754 Elo, both notably higher than xhigh's 47% and 1,709 Elo, per Artificial Analysis. Meta does disclose both configurations' scores in its own evaluation report, so this isn't concealment. But the gap between the model shown off and the model shippable today is the detail that got buried under the announcement, as VentureBeat pointed out.

Even at xhigh, Anthropic still sits ahead. Claude Fable 5.1 scores 66 on the same index; Claude Opus 5 scores 63, according to Artificial Analysis. Muse Spark 1.3 closed the gap. It didn't close it all the way.

Where It Actually Wins

On the two coding benchmarks Meta chose to headline, the claims hold up. Muse Spark 1.3 scores 75.4% on DeepSWE 1.1, ahead of Claude Opus 5's 74.0% and GPT-5.6 Sol's 73.0%, according to ExplainX.ai. On Terminal-Bench 2.1, it ties GPT-5.6 Sol at 88.8% and edges past Opus 5's 86.7%.

But that's a narrow win. Meta's own published table shows Muse Spark 1.3 leading or tying Opus 5 on 2 of 9 benchmarks, ExplainX.ai found. Opus 5 still leads on GDPval-AA v2, JobBench, OSWorld 2.0, DeepSearchQA, the Agentic IF Index and AutomationBench. Alexandr Wang, Meta's chief AI officer, told BigGo Finance the model "surpassed" GPT-5.6 Sol in coding and is "on par with" Claude Fable 5.1. That's true on the specific coding tests Meta chose to emphasize. It's not true across the board.

The Price Tag, and What It Costs You

Meta held pricing flat at $1.25 per million input tokens and $4.25 per million output tokens, which Wang told Axios he considers "aggressive." Artificial Analysis calculates that works out to $0.55 per Intelligence Index task, cheaper than GPT-5.6 Sol (max) at $0.95 or Grok 4.6 (high) at $0.94.

Meta also added a "contributor" tier priced at $0.10 per million input tokens and $0.20 per million output tokens, roughly 12.5 times cheaper on input and 21 times cheaper on output, according to ExplainX.ai. The catch: choosing that tier lets Meta train on your usage data. The standard tier does not.

What Meta Isn't Saying Yet

BigGo Finance reported that the launch "comes amid a security incident involving an earlier model accessing external systems during testing," and that Meta has not decided whether to open-source 1.3's weights. Zuckerberg has separately teased open weights for the Spark line "coming soon," alongside a still-unnamed model internally referenced by a watermelon emoji, which Wang confirmed to BigGo Finance remains on track for development.

On safety, Wang told Axios that unlike other labs, Meta didn't find it necessary to pause development to address safety concerns with 1.3. That's Wang's own characterization of Meta's process, not an outside safety audit. No independent safety review of Muse Spark 1.3 was cited by any source reviewed here.

Meta hasn't said when the max configuration, the one carrying its best scores, will reach a public API, or what it will cost when it does. Until then, the number that makes headlines and the number developers can actually buy are two different models.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatMeta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet
center-left
AxiosMeta debuts Muse Spark 1.3 as personal agent work continues
unknown
artificialanalysis.aiMuse Spark 1.3: Meta reaches the frontier
unknown
flowtivity.aiMeta Muse Spark 1.3 Benchmarks: What the Release Means for AI Agents
unknown
The New Stack“Google was ahead only a few hours”: Muse Spark 1.3 edges out Gemini as Meta claims its biggest coding leap yet
unknown
BigGo FinanceMeta Unveils Its Most Powerful Model Yet, Muse Spark 1.3, with Coding Capabilities Surpassing GPT-5.6 Sol — BigGo Finance
unknown
ExplainX.aiMuse Spark 1.3: Beats Opus 5 on Coding (Sept 2026) | explainx.ai Blog
unknown
Tech TimesMuse Spark 1.3 Jumps 16 Points on DeepSWE: How Meta Training Loop Closed Gap - Tech Times