READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Alibaba's New Qwen3.8-Max Model Nearly Matches Anthropic on Benchmarks, DeepSeek Undercuts Everyone on Price

Alibaba's New Qwen3.8-Max Model Nearly Matches Anthropic on Benchmarks, DeepSeek Undercuts Everyone on Price
Alibaba unveiled its 2.4 trillion-parameter Qwen3.8-Max on Monday, claiming performance close to Anthropic's top models. Days earlier, DeepSeek's new V4-Flash model showed up as the cheapest well-known AI model to run, per Artificial Analysis. China's AI industry is closing the gap with U.S. labs fast, and it's doing it on a fraction of the compute Washington tried to deny them.

Alibaba dropped its newest flagship AI model on Monday, and the numbers are hard to wave off.

Qwen3.8-Max runs on 2.4 trillion parameters, according to Alibaba and reporting from The Hindu. That's slightly smaller than Moonshot AI's Kimi K3, which has 2.8 trillion parameters and launched last month. But size isn't destiny in this game. Alibaba says its model uses a "mixture-of-experts" design that only activates 95 billion parameters at a time, cutting compute costs and speeding up responses.

On the crowdsourced benchmarking site Arena.AI, Qwen3.8-Max immediately became the top-ranked Chinese text model, The Hindu reported. It still trails Anthropic's Claude Fable 5 and three Opus variants. But on Arena.AI's image and visual-analysis leaderboard, it ranked second globally, beaten only by a Claude Fable 5 variant.

That's a Chinese open-weight model sitting one spot behind Anthropic on visual reasoning. Not "catching up eventually." Now.

Alibaba says the model independently ran a software engineering project for 16 days straight in internal testing, a claim reported by both The Star and The Hindu. Both models, Qwen3.8-Max and Kimi K3, can handle text, image and video input and process up to 1 million tokens at once, meaning they can chew through massive codebases or hundreds of pages of documents in a single pass.

Alibaba plans to release the Qwen3.8-Max weights for public download next week through its Model Studio platform, letting outside developers customize and deploy it themselves.

DeepSeek's angle: it's not about being the smartest, it's about being cheap

While Alibaba was chasing benchmark supremacy, DeepSeek took a different approach entirely. The startup officially released its V4-Flash model on Friday, and according to research firm Artificial Analysis, cited by Reuters, it is by far the cheapest well-known AI model to run on benchmark tests.

DeepSeek charges $0.14 per million input tokens and $0.28 per million output tokens. Artificial Analysis estimated the average cost of running V4-Flash through its test suite at 3 cents. Compare that to 86 cents for Kimi K3, $1.86 for OpenAI's GPT-5.6 Sol, and $3.15 for Anthropic's Claude Fable 5. That makes V4-Flash more than 100 times cheaper to run than Fable 5 on the same tasks, per Reuters.

Cost-per-test matters more than headline pricing because it accounts for how many steps a model needs to actually finish a job. A cheap-looking model that takes ten times longer to solve a problem isn't actually cheap.

On raw intelligence, V4-Flash scored 50 out of 100 on Artificial Analysis's Intelligence Index, tying Google's Gemini 3.6 Flash and landing one point behind Meta's Muse Spark 1.1 and Z.AI's GLM-5.2. Kimi K3 scored a 57. Anthropic's Opus 5 and Fable 5, along with OpenAI's GPT-5.6, scored nine or more points higher than V4-Flash.

So DeepSeek isn't claiming to be the smartest model on the market. It's claiming to be the cheapest one that's still good enough. That's the same playbook that made its R1 model a global sensation in early 2025, triggering a selloff in U.S. tech stocks, according to Reuters. DeepSeek is reportedly preparing for a potential IPO and is also working on a more powerful V4-Pro version, though no release date has been given.

Why this keeps happening despite export controls

Washington restricted export of advanced Nvidia chips to China specifically to slow this kind of progress. Anthropic's Fable 5 was temporarily placed under U.S. export controls because of its capabilities, according to The Star. The stated logic is straightforward: if China can't get top-tier chips, its AI models should lag behind.

Vey-Sern Ling, managing director at Union Bancaire Privée, told Bloomberg (via The Star) that "many investors continue to underestimate Chinese AI models because of US chip restrictions or general scepticism. In reality, the gap is probably much closer, and narrowing fast."

The fair counterpoint is that chip restrictions were never designed to stop China from building any AI models. They were designed to slow the pace and raise the cost. Chinese firms are still working with less raw compute than U.S. labs, and mixture-of-experts efficiency tricks are partly a response to that constraint, not proof the constraint doesn't matter. Whether that's a five-year head start evaporating or a permanent capability gap holding steady underneath the marketing is genuinely unresolved.

What's not unresolved: three major Chinese labs, Alibaba, Moonshot, and DeepSeek, plus ByteDance and Z.AI in the wings, have all released frontier-competitive models within roughly a month of each other, all open-weight, all dramatically cheaper than their American counterparts. That's the trend line that should worry anyone betting on U.S. AI dominance being permanent.

Alibaba's Qwen3.8-Max weights go public next week. DeepSeek's V4-Pro has no release date yet. Both will get tested against Anthropic and OpenAI's next moves soon enough.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

unknown
thestar.com.myAlibaba releases AI model with performance rivalling Anthropic - The Star
unknown
keloDeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says | KELO-AM
unknown
thehinduAlibaba unveils its most capable AI model to date, not far behind Moonshot's in size