READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

San Francisco Lab Poolside Releases Open-Weight Coding Model, Says It Beats Rivals 10 Times Its Size

San Francisco Lab Poolside Releases Open-Weight Coding Model, Says It Beats Rivals 10 Times Its Size
Poolside released Laguna S 2.1 on Tuesday, a 118-billion-parameter open-weight coding model the company says outperforms much larger systems from DeepSeek, Thinking Machines and Nvidia on key coding benchmarks. The bigger story is who's been winning the open-weight race: mostly Chinese labs, and Poolside is explicitly positioning itself as the Western answer.

Poolside, a San Francisco AI lab that has spent most of its three-year existence quietly selling coding models to governments and defense agencies, released its most capable model yet on Tuesday. The company is betting that transparency, not raw size, is how a smaller lab competes at the frontier of AI.

The model, called Laguna S 2.1, is a 118-billion-parameter Mixture-of-Experts system. It only activates 8 billion parameters per token, supports context windows up to 1 million tokens, and is available immediately on Hugging Face under the permissive OpenMDW-1.1 license, according to Poolside.

The Numbers Poolside Is Pointing To

Poolside reports that Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1, a benchmark measuring long-horizon terminal tasks. That places it 11th on the company's own compiled leaderboard, ahead of DeepSeek-V4-Pro-Max, a 1.6-trillion-parameter model that scored 64.0%, Thinking Machines' 975-billion-parameter Inkling at 63.8%, and Nvidia's 550-billion-parameter Nemotron 3 Ultra at 56.4%, according to VentureBeat's reporting on the company's release.

On SWE-Bench Multilingual, Poolside says the model posts 78.5%. On SWE-Bench Pro's public dataset, it scores 59.4%.

Those are Poolside's own published benchmarks. No independent third-party verification of these specific scores was cited in the release, which is standard practice across the industry but worth stating plainly: these are the lab's numbers, not an outside audit's.

Fast Turnaround

What may be more notable than any single benchmark is the timeline. Poolside began pre-training the model on May 22 and had it publicly launched in under nine weeks, trained on 4,096 Nvidia H200 GPUs, according to the company. That's the lab's third model release in three months, in an industry where flagship models typically ship on a cycle of quarters or years.

The Real Fight Is Over Who Controls Open-Weight AI

Over the past year, developers have shifted hard toward open-weight AI models, systems anyone can download, inspect, and run on their own servers instead of renting access from a closed API. That shift has been overwhelmingly won by Chinese labs. DeepSeek, Qwen, Kimi, GLM, MiniMax, and Tencent's Hunyuan all dominate the open-weight comparison tables Poolside itself uses to benchmark its new release.

Poolside's press release frames Laguna S 2.1 as a direct response to that gap, noting that no Western lab had released open weights in this size class in 11 months, since OpenAI's gpt-oss-120b last August. "The West needs open-weight models it can trust, run, and build on," Poolside co-CEO Jason Warner said in the announcement.

Co-founder and co-CEO Eiso Kant made the stakes even more explicit in a post on X. "I believe intelligence should and will become a commodity," he wrote, arguing the open ecosystem "will not win by being the best in its own category." His point: developers want the best available tool for the job, so if Western open-weight models can't match or beat their closed and Chinese-made counterparts, Western labs lose the developer base entirely.

That's a fair concern. If the default open-weight model running inside American companies, defense contractors, and government agencies comes from a lab operating under Chinese jurisdiction, that's a real supply-chain and security question, not just a marketing angle. Critics of that framing would note Poolside has an obvious commercial incentive to raise the alarm since it sells coding models to those exact same government and defense customers. Both things can be true at once.

Poolside's benchmark claims have not been independently replicated as of this writing. The company's leaderboard placements, including where DeepSeek-V4-Pro-Max, Inkling, and Nemotron 3 Ultra land relative to Laguna S 2.1, come from Poolside's own compiled comparisons rather than a neutral third party running all models under identical conditions.

The weights are public now, which means outside researchers and rival labs can run their own tests against Poolside's claims starting immediately. Given how fast this space moves, and given that Poolside has now shipped three models in three months, expect either confirmation or pushback on these numbers well before the lab's next release.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatPoolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size