READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Grok 4.5 Benchmarks Show a Capable but Not Dominant Model. The Price Argument Is Real.

Grok 4.5 Benchmarks Show a Capable but Not Dominant Model. The Price Argument Is Real.
SpaceXAI's Grok 4.5, released July 8 in partnership with Cursor, lands fourth on Artificial Analysis's agentic leaderboard but costs roughly 90% less per completed task than the models ranked above it. The cost gap is large enough that enterprise buyers may not care much about the capability gap.

Since xAI and Cursor announced their joint training partnership in April, the first concrete product of that deal, Grok 4.5, shipped Wednesday, July 8.

Where It Actually Stands Grok 4.5 is a

mixture-of-experts model trained across tens of thousands of NVIDIA GB300 GPUs, according to Cursor's official release. Training drew on trillions of tokens of Cursor user-interaction data, plus high-quality STEM research papers and broader knowledge-work material. xAI says reinforcement learning covered hundreds of thousands of tasks in realistic software engineering environments. The benchmarks tell a mixed story. On Terminal Bench 2.1, which tests complex command-line work, Grok 4.5 scores 83.3%, according to The Decoder, nearly matching OpenAI's GPT-5.5 (83.4%) and trailing Anthropic's Fable 5 (84.3%) by one point. On DeepSWE 1.1, which measures resolution of real GitHub issues, the gap widens: Grok 4.5 hits 53%, behind GPT-5.5 at 67% and Fable 5 at 70%. Artificial Analysis, an independent benchmarking firm, ranked Grok 4.5 fourth on its GDPval-AA v2 index of real-world agentic knowledge work, with an Elo score of 1543, behind Fable 5, GPT-5.5, and Opus 4.8. On the Coding Agent Index, Grok 4.5 running inside Grok Build scores 76 points, matching GPT-5.5 in Codex and trailing Fable 5 in Claude Code by a single point. Elon Musk didn't try to spin that. "Our internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster," he wrote on X. "We are closing the loop on real-world usefulness, not benchmarks. Hardcore engineers at Tesla and SpaceX find Grok 4.5 genuinely useful, which is what actually matters."

The Price Case Grok 4.5 is

priced at $2 per million input tokens and $6 per million output tokens. Anthropic's Fable 5 runs $10 input and $50 output. GPT-5.5 and GPT-5.6 sit at $5 input and $30 output. Opus 4.8 charges $5 input and $25 output, according to The Decoder. On per-task cost for agentic workloads, Artificial Analysis measured Grok 4.5 at $0.49 per completed task, nearly 90% cheaper than the models ranked above it, according to VentureBeat. In Grok Build specifically, the cost comes in at $2.49 per task, versus $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code, per The Decoder's figures. Grok 4.5 also uses far fewer tokens per task: an average of 1.9 million, compared to 6.2 million for GPT-5.5 and 7.2 million for Fable 5. For enterprise teams running agents across hundreds of developers for hours at a stretch, token efficiency matters significantly. Investor Gavin Baker put it plainly on X: "Pareto dominant for coding by the numbers. We will see on the all-important vibes."

The Strongest Case Against

Skeptics have a legitimate concern, and it's not the capability gap. Artificial Analysis flagged a significant problem: Grok 4.5's accuracy on the AA-Omniscience Index rose from 35% to 52%, but its hallucination rate jumped from 25% to54%, according to The Decoder. The model knows more, and it is more confidently wrong more often. For coding agents running autonomously across long sessions, a high hallucination rate poses real risk. An agent that confidently produces plausible-but-wrong code and doesn't flag the error can create compounding problems that take hours to diagnose. xAI has not publicly addressed this data point. Cursor's release notes mention "new safeguards reflecting the model's cybersecurity capabilities" but don't quantify them.

What Cursor's Role Actually Means

This model represents the first fruits of a deal that is still structurally unresolved. xAI and Cursor struck a partnership in April that could end either in xAI investing $10 billion into Cursor or acquiring it outright for $60 billion, according to Engadget. No decision has been announced. For now, Grok 4.5 is available across Cursor's desktop, web, iOS, CLI, and SDK. Individual and team plan subscribers get doubled usage for the first week. Cursor's release notes also clarify that Grok 4.5 and its previous Composer 2.5 model are different weight classes and will both remain available. The company says it intends to keep releasing new models in the Composer size range, meaning Grok 4.5 does not replace the smaller model, it sits above it. One firm geographic limit: Grok 4.5 is NOT available in the European Union as of July 9. xAI expects EU availability in mid-July, according to Engadget, without specifying an exact date. The unresolved question sitting over all of this is whether Cursor's acqui-hire or acquisition completes before the next model generation ships. If xAI closes the $60 billion deal, the joint training arrangement becomes internal, and future models would be built under a single corporate roof. That changes the competitive dynamics considerably for Anthropic and OpenAI, neither of whom owns a major developer-facing coding platform.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatSpaceX's Grok 4.5 launches at half the price of rivals — here's why that could rattle Anthropic and OpenAI | VentureBeat
center-left
EngadgetSpaceXAI launches Grok 4.5, its first built with Cursor's help
unknown
cursorIntroducing Grok 4.5 - Cursor
unknown
the-decoderGrok 4.5 is so cheap compared to Fable 5 and GPT 5.5 that benchmark gaps may not matter much - The Decoder