READ. SCROLL. LISTEN.

Unbiased headlines. Facts, not spin.

Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Elon Musk's SpaceXAI Releases Grok 4.6, Matches GPT-5.6 Sol on Independent Benchmark

Elon Musk's SpaceXAI Releases Grok 4.6, Matches GPT-5.6 Sol on Independent Benchmark
SpaceXAI, formerly xAI, released Grok 4.6 on August 12, scoring 61 on the Artificial Analysis Intelligence Index, tying OpenAI's GPT-5.6 Sol and beating China's Kimi K3. The model keeps Grok 4.5's pricing while posting real agentic gains, which matters more than the leaderboard score.

Elon Musk's xAI released its Grok 4.6 model on August 12. It scored 61 on the Artificial Analysis Intelligence Index, a third-party benchmark that tracks frontier AI models.

That score ties OpenAI's GPT-5.6 Sol at its maximum reasoning setting, according to Artificial Analysis. It beats Moonshot AI's popular open-weight Chinese model Kimi K3, and lands five points above xAI's own Grok 4.5, released just over a month earlier. Claude Opus 5 and Claude Fable 5, both from Anthropic, still sit ahead of everyone at 63 and 62.

The headline number is a leaderboard curiosity. The pricing decision is the actual story.

xAI held Grok 4.6's API pricing flat at $2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens, the same rate as Grok 4.5, according to xAI's own release notes. Above 200,000 tokens, pricing rises to $4 and $12. Compare that to Claude Opus 5 at $5/$25 and GPT-5.6 Sol's standard mode at $5/$30, both more than four times Grok's cost on the output side.

Artificial Analysis called flat pricing across a model generation "unusual at the frontier," noting that intelligence gains have typically come with price increases, not stable ones. Grok 4.6 cost $0.84 per task in Artificial Analysis's testing, matching Kimi K3's cost efficiency while scoring higher on the intelligence index.

Where Grok 4.6 actually wins

The benchmark tie with GPT-5.6 Sol undersells what's different about this release. xAI built Grok 4.6 specifically for long-running agent work, coding, and knowledge tasks that require a model to stay on track across many steps without losing the thread.

On GDPval-AA v2, a benchmark for real-world agentic knowledge work, Grok 4.6 posted an Elo of 1753, trailing only Claude Opus 5 and statistically tied with Claude Fable 5 and Alibaba's Qwen3.8 Max, according to Artificial Analysis. On Terminal-Bench v2.1, which tests terminal-based software tasks, it scored 88.4%, putting it level with the leading models.

The turn efficiency numbers stand out most. Artificial Analysis found Grok 4.6 completes agentic tasks in about 53 turns using roughly 0.5 billion input tokens on average. Claude Opus 5 needs about 103 turns and 2.0 billion input tokens for comparable work. Fewer turns and fewer tokens means lower real-world cost for companies running these models at scale, on top of the already lower per-token price.

The Cursor and Grok Bot push

Grok 4.6 didn't launch in isolation. xAI acquired the AI coding startup Cursor in June, and the two companies released a new agent product called Grok Bot on August 11, one day before Grok 4.6's launch, according to 9to5Mac. Grok Bot gives users a persistent cloud-based AI agent with messaging, approval workflows, and automated routines.

Grok 4.6 is available now through Cursor, xAI's own Grok Build coding tool, the xAI API, and third-party platforms including OpenRouter, Vercel, and Cloudflare. Grok Build access starts at $30 a month through the SuperGrok plan. xAI is offering double the usual usage limits for Grok 4.6 in both Cursor and Grok Build during its first week on the market.

9to5Mac framed the release as xAI "repairing Grok's reputation," pointing to the Cursor acquisition and Grok Bot launch as part of a broader push to be taken seriously in coding and enterprise agent work, not just as a chatbot competing with ChatGPT.

What's still unresolved

Artificial Analysis's numbers come from a third-party lab that runs its own standardized tests, not from xAI's marketing claims, which gives the benchmark more credibility than a company's self-reported scores. But independent, real-world enterprise adoption data for Grok 4.6 doesn't exist yet, since the model has been public for less than a day as of this writing.

Whether Grok 4.6's turn-efficiency edge translates into actual cost savings for companies running large coding or agent workloads will depend on how the model performs outside of controlled benchmarks. That's the test that actually matters, and it hasn't happened yet.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatSpaceXAI debuts Grok 4.6, overtaking Kimi K3's performance and matching GPT-5.6 Sol for world's third best on Artificial Analysis
unknown
docs.x.aiRelease Notes | SpaceXAI Docs
unknown
9to5macSpaceXAI releases Grok 4.6, claiming GPT-5.6 Sol and Claude Fable 5-level intelligence
unknown
artificialanalysis.aiGrok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency