Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Meta's Muse Spark 1.3 Ties OpenAI and Grok on Benchmarks, But Its Best Score Comes From a Version Nobody Can Fully Access Yet

Since Meta's first Muse Spark model launched April 8, 2026, the company has pushed out three follow-on versions in five months. Muse Spark 1.3 rolled into Meta's Muse Code terminal agent and the Meta Model API on Wednesday, September 2, with broader paid API access opening today, according to flowtivity.ai and BigGo Finance. It's the fourth release in the line and the one Meta is calling its biggest jump yet.
Mark Zuckerberg announced it on X, writing that Muse Spark 1.3 delivers "frontier performance almost too cheap to meter" and calling it Meta's "biggest jump we've made so far on coding and agentic work," according to VentureBeat and ExplainX.ai. That's a strong claim from the CEO. The numbers back some of it up, and complicate the rest.
Two Versions, One Headline Number
Meta shipped two configurations of 1.3. The one developers can actually use, called xhigh, scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol (max) and Grok 4.6 (high), according to Artificial Analysis, the independent benchmarking firm. That's up from 57 for Muse Spark 1.2 in August and 51 for 1.1 in July.
The other configuration, called max, scores 62, according to Artificial Analysis. But max is still in limited partner preview and completing safety testing, VentureBeat reported. Artificial Analysis says it evaluated max only in that limited preview and currently lists no API provider offering it at all.
Meta's launch materials prominently feature max's results, including a #1 score on Tau3-Bench Banking at 52% and a GDPval-AA v2 score of 1,754 Elo, both notably higher than xhigh's 47% and 1,709 Elo, per Artificial Analysis. Meta does disclose both configurations' scores in its own evaluation report, so this isn't concealment. But the gap between the model shown off and the model shippable today is the detail that got buried under the announcement, as VentureBeat pointed out.
Even at xhigh, Anthropic still sits ahead. Claude Fable 5.1 scores 66 on the same index; Claude Opus 5 scores 63, according to Artificial Analysis. Muse Spark 1.3 closed the gap. It didn't close it all the way.
Where It Actually Wins
On the two coding benchmarks Meta chose to headline, the claims hold up. Muse Spark 1.3 scores 75.4% on DeepSWE 1.1, ahead of Claude Opus 5's 74.0% and GPT-5.6 Sol's 73.0%, according to ExplainX.ai. On Terminal-Bench 2.1, it ties GPT-5.6 Sol at 88.8% and edges past Opus 5's 86.7%.
But that's a narrow win. Meta's own published table shows Muse Spark 1.3 leading or tying Opus 5 on 2 of 9 benchmarks, ExplainX.ai found. Opus 5 still leads on GDPval-AA v2, JobBench, OSWorld 2.0, DeepSearchQA, the Agentic IF Index and AutomationBench. Alexandr Wang, Meta's chief AI officer, told BigGo Finance the model "surpassed" GPT-5.6 Sol in coding and is "on par with" Claude Fable 5.1. That's true on the specific coding tests Meta chose to emphasize. It's not true across the board.
The Price Tag, and What It Costs You
Meta held pricing flat at $1.25 per million input tokens and $4.25 per million output tokens, which Wang told Axios he considers "aggressive." Artificial Analysis calculates that works out to $0.55 per Intelligence Index task, cheaper than GPT-5.6 Sol (max) at $0.95 or Grok 4.6 (high) at $0.94.
Meta also added a "contributor" tier priced at $0.10 per million input tokens and $0.20 per million output tokens, roughly 12.5 times cheaper on input and 21 times cheaper on output, according to ExplainX.ai. The catch: choosing that tier lets Meta train on your usage data. The standard tier does not.
What Meta Isn't Saying Yet
BigGo Finance reported that the launch "comes amid a security incident involving an earlier model accessing external systems during testing," and that Meta has not decided whether to open-source 1.3's weights. Zuckerberg has separately teased open weights for the Spark line "coming soon," alongside a still-unnamed model internally referenced by a watermelon emoji, which Wang confirmed to BigGo Finance remains on track for development.
On safety, Wang told Axios that unlike other labs, Meta didn't find it necessary to pause development to address safety concerns with 1.3. That's Wang's own characterization of Meta's process, not an outside safety audit. No independent safety review of Muse Spark 1.3 was cited by any source reviewed here.
Meta hasn't said when the max configuration, the one carrying its best scores, will reach a public API, or what it will cost when it does. Until then, the number that makes headlines and the number developers can actually buy are two different models.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.