Original briefings. Zero spin.
Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.
Mira Murati's Thinking Machines Releases Inkling-Small, a Smaller Open-Source AI Model That Nearly Matches Its Flagship

Thinking Machines Lab released Inkling-Small, a scaled-down version of its first open-weight AI model, according to VentureBeat and the company's own announcement on its website. The release comes about two weeks after the startup, led by former OpenAI chief technology officer Mira Murati, put out its original model, Inkling.
The numbers are the story here. Inkling, the flagship, has 975 billion total parameters and 41 billion active parameters per token, according to Thinking Machines. Inkling-Small runs at 276 billion total parameters with just 12 billion active per token, roughly a quarter the active compute load. Despite that shrinkage, Inkling-Small scored 40 on the Artificial Analysis Intelligence Index, a third-party benchmark, compared with 41 for the full-size Inkling, VentureBeat reported.
Artificial Analysis also said no open-weight model at Inkling-Small's size or smaller scored higher on that index, according to VentureBeat.
Where the small model actually wins
Benchmark parity alone would be notable. But Thinking Machines and VentureBeat both report Inkling-Small outright beating its bigger sibling on specific tests: 80.2% versus 77.6% on SWE-bench Verified, a coding benchmark, and 64.7% versus 63.8% on Terminal Bench 2.1. VentureBeat also reported Inkling-Small edging ahead on SciCode, Humanity's Last Exam, GPQA Diamond, and CritPt.
The gains aren't universal. Inkling still holds a clear lead on factual-knowledge tasks and some agentic benchmarks. VentureBeat cited one figure showing Inkling-Small scoring 15.5% on a banking-agent task called τ3-Banking, well below Inkling's result on the same measure, though the exact flagship score wasn't specified in available reporting.
What it's built for
Both models handle text, image, and audio input and produce text output, with a context window up to one million tokens, according to Thinking Machines' own release notes. Inkling was pretrained on 45 trillion tokens spanning text, images, audio and video, the company said.
Thinking Machines describes Inkling not as the strongest model available, open or closed, but as a broad, flexible base meant for fine-tuning. That's a notably modest claim for a company that just raised eye-popping funding and hired Murati away from OpenAI. The company is betting on customization, not raw benchmark supremacy, as the pitch to enterprise buyers.
To back that pitch, Thinking Machines is offering full model weights on Hugging Face, a fine-tuning API called Tinker, and a new "Inkling Playground" chat interface inside the Tinker console for developers to test-drive the model before committing engineering time to it.
The price cut
For companies actually trying to deploy this, cost matters more than benchmark decimals. Thinking Machines is running a launch discount of 50% off, according to VentureBeat, bringing the standard 64K-context version of Inkling-Small to $0.58 per million input tokens, $1.44 per million output tokens, and $1.73 per million training tokens. Cached input requests run cheaper still, at $0.116 per million tokens. A larger 256K-context version costs more.
That pricing structure, combined with the smaller active-parameter count, is the actual sales pitch. Running a 975-billion-parameter model requires serious GPU infrastructure that most companies don't have and can't easily rent at scale. A model that needs a quarter of the active compute but gives up almost nothing on real-world coding and reasoning tasks is a meaningfully different proposition for a mid-size firm evaluating whether to build on open weights instead of paying for a closed API from OpenAI, Google, or Anthropic.
Still, Inkling-Small is not something you run on a laptop. VentureBeat is direct about this: the model remains "far too large for a laptop or conventional workstation." It's aimed at enterprises with some in-house GPU capacity, not solo developers or small shops.
What's unresolved
Thinking Machines describes Inkling-Small as a "preview" release in its own announcement, and calls Inkling "the first in a family of models of different sizes." That leaves open how many more variants are coming, on what timeline, and whether the company's promised model family will include something small enough for on-device or consumer use.
These benchmark comparisons come largely from Thinking Machines' own released figures and a single third-party evaluator, Artificial Analysis. Independent, adversarial testing by enterprises actually deploying Inkling-Small at scale, particularly on agentic and tool-use tasks where the model reportedly lags, hasn't been reported yet.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.