Original briefings. Zero spin.
Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.
Open-Weight AI Models Are Eating into Frontier Lab Margins, and the Industry Knows It

The threat is not just falling token prices in the abstract. It is open-weight models becoming capable enough to handle most enterprise workloads without touching a frontier API at all.
The Scorecard Is Changing
For two years, the AI race ran on one metric: biggest model, best benchmark. That is no longer how corporate buyers are shopping.
Perplexity CEO Aravind Srinivas told CNBC the product is no longer the model itself. "It is the harness, the orchestration system that puts the model inside a very capable harness and pairs the model with a lot of tools."
What that means in practice: companies are building systems that route tasks to whichever model is cheapest for that specific job, escalating to a more powerful model only when the task demands it. Customer service query? Run it on something cheap. Complex code review? Call in the bigger model. Routine internal workflow? Open-weight, self-hosted, no API bill. "The answer is always use whatever is the best for the task," Srinivas said.
This shift is emerging as corporate America tightens its belt on AI spending, presenting another challenge for OpenAI and Anthropic, which have flourished by selling the most cutting-edge technology.
Perplexity's Move Signals the Direction
This week, Perplexity previewed a computer-use product built around GLM 5.2, an open model from China's Z.ai. The system is explicitly designed to let a cheaper model handle the bulk of the work while reserving stronger models for harder steps. According to CNBC, that approach reflects a market-wide shift, not a Perplexity-specific experiment.
Benchmark's Call: 90% of Tokens Go Open-Weight
Benchmark general partner Peter Fenton put a number on it. "A maybe contrarian view that is becoming consensus is our belief that 90-plus percent of the tokens created will come out of open-weight models over the next 18 to 24 months, possibly even by the end of the year," Fenton told CNBC.
On margins: "The inference margins generated by the frontier model companies, I think, are going to come under pressure when you can run those without the markup that they're providing, when you have good enough models from open weights."
Fenton's firm invested in Ollama, a company that makes it easier for developers and enterprises to download, run and manage open models — so he is an interested source. But the underlying logic tracks with what enterprise buyers are already doing: tightening AI budgets and demanding cost justification.
Fenton also noted the move to open models is not only about saving money. In some cases, smaller models tuned for a specific task can be faster and perform better than larger general-purpose models.
Adoption Is Already Broad
Ollama CEO Jeff Morgan said Ollama has been adopted by more than 85% of the Fortune 500, including companies in regulated industries such as aviation, insurance and health care. He said many companies start with smaller models running close to their own data, then expand to larger open models as they get more comfortable.
"One thing is where the model's from and where it was created and trained," Morgan said. "But the more important thing to these businesses we speak to is where it runs and how it runs."
The Strongest Counter-Argument
There is a real case for the frontier labs holding their position. Open-weight models require companies to run their own infrastructure, hire engineers to tune and maintain them, and absorb the security and compliance burden that comes with self-hosting. For most mid-sized enterprises, that total cost of ownership is not trivial. OpenAI and Anthropic are not just selling model quality. They are selling a managed, maintained, legally accountable service. A Fortune 500 company with regulatory exposure may gladly pay a premium to not own that problem.
The frontier labs are also not standing still. They continue to push capability ceilings that open-weight models chase but have not matched on every task. For genuinely hard, novel reasoning tasks, the gap remains real.
Still, "most enterprise workloads" are not genuinely hard reasoning tasks. They are structured, repetitive, and well-suited to a tuned smaller model. That is where Fenton's 90% figure gets traction.
The Datacenter Question
The current AI boom assumes demand will keep flowing to large cloud data centers filled with high-end chips. Srinivas says some AI work may eventually run locally instead, on devices owned by consumers or businesses. That would not eliminate the need for data centers, but it could create a more hybrid AI system, with routine tasks run locally and the most difficult work sent to a more powerful model in the cloud.
If Fenton is right — if 90% of token volume migrates to open-weight models that companies run themselves — the companies that built data centers to serve frontier API inference do not collect revenue on most of those tokens. The volume exists; the billing relationship does not.
The China Dimension
GLM 5.2 from Z.ai is a Chinese model. Perplexity's decision to build a product around it — specifically because it is capable enough and cheap enough — is a data point on where open-weight competition is coming from. The rise of competitive open-weight models from Chinese labs, including Z.ai and DeepSeek, has made open-source AI a business issue, a policy issue and a national competitiveness issue, according to CNBC.
Srinivas argued the U.S. should support open models because they make AI more affordable and accessible. "If you want the benefits of AI to be widely distributed to small businesses in America and American allied countries, then you really need AI to be a lot more affordable," Srinivas said. "And open source is the only way to do that."
For investors, the question is whether the biggest AI labs can maintain their pricing power as open models get better and companies become more selective about what they use.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.