READ. SCROLL. LISTEN.

Unbiased headlines. Facts, not spin.

Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

AI Price War: OpenAI and Anthropic Slash Costs as Chinese Rivals Undercut the Market

AI Price War: OpenAI and Anthropic Slash Costs as Chinese Rivals Undercut the Market
OpenAI cut prices on two GPT-5.6 tiers by up to 80% at the end of July, and Anthropic launched Claude Opus 5 at half its predecessor's price, both responding to cheap Chinese models from DeepSeek and Moonshot AI. Usage is rising faster than prices are falling, but CFO surveys show many enterprises are still blowing through AI budgets because total consumption keeps climbing.

Running AI models just got a lot cheaper, and it happened fast.

Average enterprise inference prices hit a 2026 low of $1.16 to $1.18 per million tokens between August 6 and 8, according to Jefferies analysts citing data from Silicon Data, as reported by the South China Morning Post. That's down from $2.04 on May 31 and $1.45 in late July. Silicon Data's index tracks pricing across API providers and open-weight platforms used by developers worldwide.

The trigger is straightforward: Chinese AI labs are shipping models that come close to matching US performance at a fraction of the cost. DeepSeek and Moonshot AI, maker of Kimi K3, are the two names driving the pressure, according to Tom's Hardware and Crypto Briefing. NTD reported that DeepSeek V4 Flash approaches Claude Opus-level performance on coding and reasoning tasks while charging around $0.28 per million output tokens, compared to $25 to $50 for top-tier US models.

American labs blinked first on pricing, not performance.

OpenAI cut the price of GPT-5.6 Luna, its fast, low-to-mid tier model, by 80% on July 30, according to Crypto Briefing and Tom's Hardware. Input pricing dropped from $1.00 to $0.20 per million tokens; output fell from $6.00 to $1.20. The mid-tier Terra model got a 20% cut. OpenAI's flagship Sol model held its price, though the company sped up its top-tier performance options at the same cost, a signal it's protecting margin on its best product while fighting for volume everywhere else.

Anthropic followed with Claude Opus 5, priced at roughly half of its predecessor Fable 5, at $5 per million input tokens and $25 per million output tokens, according to NTD. Anthropic also canceled a planned price hike on its Sonnet 5 model. Google, meanwhile, rolled out cheaper Gemini 3.6 Flash and 3.5 Flash-Lite models about a week before OpenAI's cuts, according to Tom's Hardware.

The speed of the collapse stands out in the numbers. Tom's Hardware points out that OpenAI launched GPT-5.4 in March at $2.50 input and $15 output per million tokens. Five months later, Luna is priced at $0.20 and $1.20. A frontier-priced model got discounted into budget territory in under four months. Crypto Briefing frames the longer arc: GPT-4-class performance cost over $20 per million tokens in late 2022. That same capability tier now costs less than $1 across multiple providers, a drop of more than 95% in roughly three and a half years.

Usage isn't falling with prices. It's exploding.

TD Cowen analyzed usage data from OpenRouter, a marketplace that lets developers access various models, and found that after Luna's price cut, its effective price fell roughly tenfold while consumption jumped about 14-fold, according to Business Insider. Terra's effective price fell threefold while usage rose about fivefold. The result: OpenAI's revenue from Luna rose an estimated 34% compared to the week before the cut, and Terra revenue rose about 45%. Cutting a product's price 80% and making more money off it is not how pricing normally works. It's the pattern economists call Jevons Paradox, first observed in the 19th century when cheaper coal use drove total coal consumption up, not down.

That's the optimistic read. There's a real counter-argument, and it's not coming from AI skeptics grasping at straws.

McKinsey found that 93% of surveyed firms already exceed their AI budgets, according to CFO Dive. Mavvrik's research found 43% of companies cite token costs as their top source of unexpected AI spending, and only 11% of organizations can forecast their AI spending within plus or minus 10%, down from 15% in 2025. Gartner projects worldwide AI spending will hit $2.59 trillion in 2026, a 47% jump, with infrastructure eating the largest share, according to analyst John-David Lovelock cited by AI Certs. The concern from CFOs is legitimate: per-token discounts don't matter if total consumption outpaces them, and "shadow IT" spending on AI coding tools, used by 98% of engineering organizations but tracked in cost reporting by only 42% of them, makes the problem harder to see coming.

Both things can be true at once. Unit prices are falling and enterprises are still spending more overall, because the volume of AI work companies are doing is scaling faster than the discounts. That's not a contradiction, it's Jevons Paradox playing out on corporate balance sheets in real time.

The unresolved question is what this does to AI lab valuations built on the assumption that inference stays a high-margin business. If Chinese open-source models keep closing the capability gap while undercutting price by 10x to 50x, as NTD reported, US labs face a structural choice: keep matching Chinese prices and hope volume growth like Luna's 14-fold usage spike offsets thinner per-query margins, or cede the budget tier entirely and compete only at the frontier, where Sol-class pricing has so far held firm.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
Crypto BriefingUS labs cut AI inference costs nearly 25% amid price war
center-left
SCMPEnterprise AI costs hit 2026 low driven by price wars, Chinese open-source models: research | South China Morning Post
center-left
Business InsiderAI prices are plunging. That feels bad, but something else is happening.
unknown
cfodive1 in 4 companies delay or cancel an AI project over cost
unknown
ntdUS AI Leaders Pivot to Price Cuts as Low-Cost Chinese Rivals Flood the Market
unknown
aicerts.aiEnterprise AI Inference Costs Defy 2026 “Low” Narrative
unknown
tomshardwareAI companies are now racing to the bottom — crashing token prices and competitive models push companies to cut costs