READ. SCROLL. LISTEN.

Unbiased headlines. Facts, not spin.

Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

China's AI Price War Splits in Two: Rock-Bottom Models Get Cheaper, Frontier Models Get Pricier

China's AI Price War Splits in Two: Rock-Bottom Models Get Cheaper, Frontier Models Get Pricier
China's AI companies are no longer just racing to zero. DeepSeek launched a sub-penny model this month while jacking up prices on its flagship V4-Pro system by as much as 1,100% in August. Bank of America says the real fight now is cost-per-completed-task, not cost-per-token, and the infrastructure layer, not the model makers, is positioned to win either way.

China's AI price war just got more complicated. Instead of one race to the bottom, there are now two tracks: dirt-cheap commodity models getting cheaper, and frontier models getting more expensive.

DeepSeek launched V4.1 Flash on September 10 at under one cent per million tokens, according to BigGo Finance, and announced it would retire its V4-Pro model starting September 14, automatically routing users to the cheaper Flash version. The same week, Zhipu's overseas model Ox Alpha (GLM-5.3-Flash) ended its half-price launch promotion on September 9, with input pricing doubling back to $0.15 per million tokens. BigGo Finance reported the moves triggered sell-offs in Hong Kong-listed AI stocks and South Korean memory chip shares, including Samsung Electronics and SK Hynix.

That's the cheap end. On the expensive end, something different happened in August. DeepSeek warned customers to brace for a "significant increase" in pricing and then rolled out new peak and off-peak rates for its V4-Pro and V4-Flash models, with some charges multiple times higher than before, according to Bloomberg's Catherine Thorbecke, writing in The Japan Times. Business Insider, citing a Bank of America report from analyst Winnie Wu, put a number on that hike: DeepSeek raised V4 prices anywhere from 50% to 1,100%. OpenAI, meanwhile, cut its own Luna model prices roughly 80% in August, according to State Street's Ninghui Liu and Alvin Lau.

Bank of America's take, per Business Insider: the price war isn't ending, it's evolving. "China's LLM price war is evolving from blanket cuts to tiered pricing," Wu's team wrote. "Basic inference is increasingly commoditized, while frontier capabilities remain differentiated." The metric that matters now is cost per completed task, not price per token.

A pipeline, not a one-off

DeepSeek's original R1 model shocked global markets in January 2025 by proving a Chinese lab could match Silicon Valley for a fraction of the cost. What's happening now looks different.

Across roughly eight weeks in mid-2026, five Chinese developers shipped six frontier-adjacent releases, according to State Street's research: Moonshot's Kimi K3, Z.ai's GLM-5.2, DeepSeek's V4-Flash and V4-Pro, Alibaba's Qwen3.8-Max, and ByteDance's Seedance 2.5. State Street calls this "an industrialized pipeline" rather than a repeat of the DeepSeek shock, noting that Chinese open-weight models now top global usage charts on the developer platform OpenRouter.

But usage isn't the same as revenue. According to Singapore's Business Times, OpenRouter data through August shows eight of the ten highest-volume large language models by tokens processed were Chinese, while only two were American. Yet US models still capture more overall spending. Across nearly 30 task categories tracked, Chinese models led in only seven; ChatGPT and Claude dominated most of the rest. Kendra Schaefer of Trivium China told the Business Times that companies increasingly mix and match models by task, cost, and data security needs. "There's not one market for AI models; there's multiple market segments," she said.

Real-world users illustrate the split. Los Angeles-based entrepreneur Li Jianian told the Business Times he now runs Zhipu's GLM-5.2 on a local server for about 70% of the coding work he used to hand to Anthropic's Claude, calling the performance "already on a par with Claude" for most tasks. Zhipu's GLM-5.2 costs $4.40 per million output tokens versus $10 for Claude Sonnet 5. Meanwhile a Chengdu-based video creator identified only as Cheng Xiao still relies on a VPN to reach Claude for his hardest coding jobs, since Anthropic doesn't operate in China, and switched to OpenAI's Codex when his account got blocked again.

Who actually profits

State Street's report argues the durable winners aren't the model developers at all. Chips, semiconductor equipment, memory, and large cloud platforms have the strongest long-term economics, the researchers wrote, because compute stays scarce even as the price of intelligence keeps falling. Hyperscalers "gain twice," from rising demand and from billing usage on the open models they host. Model developers, by contrast, face "the biggest challenge," since fast-rotating leadership makes it hard to justify premium prices for very long. The LA Times noted that Moonshot's Kimi K3 ranked third globally in model intelligence in July, according to benchmarking platform Artificial Analysis, and has since slid to ninth behind newer releases from Anthropic, OpenAI, Meta and xAI.

That volatility cuts against a simple "China is winning" narrative. Samm Sacks of Johns Hopkins' Institute for America, China, and the Future of Global Affairs told the LA Times that "the gap between U.S. and Chinese models is narrow and fragile." State Street's researchers flag four things that could stall China's momentum: a durable capability lead from closed US labs, tighter compute restrictions, hidden costs of running open-weight models once hosting and security are factored in, and a broader cooling of AI spending industry-wide. None of those has happened yet, but all are live possibilities.

There's also an irony underneath it all. Years of US-led export controls on advanced chips were designed to slow China's AI progress. Business Insider notes those restrictions instead expanded the market for domestic Chinese chip and cloud suppliers, intensifying competition inside China rather than choking it off. Whether that outcome reflects a policy that backfired or one that simply hasn't finished playing out is an open question.

For now, the practical fight is happening in pricing tables. DeepSeek is charging more for its top-shelf model while giving away a near-free one. Zhipu's API business already makes up 86.5% of its first-half revenue, per BigGo Finance, a sign the company is betting its future on volume, not margin. Whether that bet pays off remains unanswered. As Bloomberg's Thorbecke asked in The Japan Times, "how much longer can China afford cheap AI?" Xi Jinping is expected to visit Washington this month to meet President Trump, according to the LA Times, putting the AI rivalry squarely on the agenda at the same moment the pricing math on both sides is still shifting.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
The Japan TimesHow much longer can China afford cheap AI?
center-left
Business InsiderChina's AI price war is entering a new phase
center-left
LA TimesAs the world debates the risks of AI, China is catching up
unknown
ssgaBeyond DeepSeek: China's 2026 model wave and the repricing of the AI stack
unknown
BigGo FinanceChina's AI Model Price War Goes Global: Zhipu's Promo Ends, DeepSeek Unleashes Another Ultra-Low-Price Blow — BigGo Finance
unknown
The Business TimesChina’s AI catch-up is forcing Silicon Valley to cut prices