READ. SCROLL. LISTEN.

Unbiased headlines. Facts, not spin.

Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Chinese AI Firm Z.ai Says Its Model Beats Anthropic on Finding Bugs, Loses on Exploiting Them

Chinese AI Firm Z.ai Says Its Model Beats Anthropic on Finding Bugs, Loses on Exploiting Them
Z.ai's new GLM-5.3 model reportedly edged out Anthropic's Mythos 5 and OpenAI's GPT-5.6 on a benchmark for spotting security flaws, but fell well behind both on actually turning those flaws into working attacks. The numbers come entirely from Z.ai itself and have not been independently verified.

Chinese AI startup Z.ai says its newest model can find software vulnerabilities almost as well as Anthropic's most locked-down cybersecurity system, and slightly better on one specific test. It's also a claim made entirely by the company selling the product.

Z.ai, also known as Zhipu, announced Friday that its open-source GLM-5.3 model scored 84.5% on CyberGym, a benchmark that tests whether an AI can review code, flag security flaws, and confirm those flaws are real, according to Reuters. That edges out the 83.8% Z.ai says Anthropic's Mythos 5 scored, and the 83.6% it says OpenAI's GPT-5.6 Sol scored, according to the South China Morning Post.

Here's the catch: none of these numbers have been independently verified. They come from Z.ai's own testing. Reuters flagged this directly. So did every other outlet covering the story. Take the win-over-America headline with a grain of salt until outside researchers replicate it.

Where Z.ai actually lost, and lost badly

Finding a bug is one thing. Turning it into a working attack is another. That's what ExploitBench measures, and it's where the gap opens up fast.

GLM-5.3 scored 54.4% on ExploitBench, according to Reuters and the South China Morning Post. Mythos 5 scored 78%. GPT-5.6 Sol scored 76.5%. That's not a close race. That's a 20-plus point gap on the part of the test that measures actual offensive capability.

In a timed test, GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours. Mythos 5 completed 181 and 247 in the same windows, according to Reuters. On raw exploit development speed and volume, the American model isn't just ahead—it's not particularly close.

For defenders, this matters. It tells you what GLM-5.3 is good at right now: reading code and spotting problems. It is not yet as good at weaponizing those problems into working attacks. For anyone worried about offensive capability slipping into more hands, that's the more important number in the whole release.

Why Anthropic locks its version down

Anthropic built Mythos as a stripped-down version of its Claude Fable 5 model, with cybersecurity safeguards removed specifically so it can hunt for vulnerabilities aggressively. It only hands that version to vetted organizations, according to Reuters. The logic is straightforward: a model good at finding and exploiting flaws helps defenders patch systems, but the same capability helps attackers break in. Anthropic isn't handing that kind of tool to anyone with an internet connection.

Z.ai says it will do something similar. The company plans to release GLM-5.3 publicly in about two weeks, after finishing security assessments, according to Reuters. Its most sensitive cyber functions will reportedly go only to verified users through what it calls a "trusted access" programme. Z.ai even echoed Anthropic's language on its own limited-access rollout, referencing Anthropic's "Project Glasswing" scheme in a Friday post on X, according to Reuters and Channel News Asia.

Gabriel Wagner, an AI governance researcher at Concordia AI, a Beijing-based AI safety consultancy, called this a notable shift. "To the best of my knowledge, this is the first time a Chinese lab is publicly justifying a delayed open release of model weights with safety considerations," Wagner told Reuters. He added that it shows "open-weight risk management practices in China are becoming more sophisticated."

That's a fair point, and it deserves to be taken at face value rather than dismissed as PR. If Chinese labs are starting to build in the same kind of staged, vetted-access release process that Western labs use, that's a real development in how open-source AI gets handled globally, not just a marketing line.

The obvious problem with any of this

Once a model's weights are actually out in the open, restrictions get harder to enforce. Critics quoted by Reuters and Channel News Asia made this point plainly: safeguards baked into a model become far weaker once outside developers can download it, modify it, or bolt it onto other tools. Z.ai says it has added layers meant to screen risky requests, monitor the model's activity, and train it to refuse malicious tasks. Whether those hold up once GLM-5.3 is in the wild is untested.

Z.ai also told The News International that GLM-5.3's cybersecurity abilities emerged from post-training on a general-purpose coding model, rather than being purpose-built as a security tool. This means these capabilities may be somewhat incidental—a byproduct of building a strong coding assistant, not a dedicated cyberweapon. It also means similar capability could show up in other general coding models going forward, whether their makers intend it or not.

Zhipu told the South China Morning Post it tested the model against 269 real-world codebases with Chinese security teams, turning up 2,436 vulnerabilities, of which 1,097 were rated medium-to-high severity after expert review. That's a genuinely large-scale test. But it's also Zhipu grading its own homework. Independent verification of any of these figures—from CyberGym to ExploitBench to the vulnerability count—has not yet happened, and until it does, this is a company's press release dressed up as a benchmark result.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
enterpriseai.economictimes.indiatimesMaking AI Work: China's Z.ai says new model nears Anthropic's Mythos 5 in cyber-defence tests, ETEnterpriseai
center-left
SCMPZhipu launches flagship model GLM-5.3 as China seeks Mythos-level edge in cyber defence | South China Morning Post
center-right
The News InternationalChina’s Z.ai unveils new AI model nearing Anthropic’s Mythos 5 in cyber tests | Technology | thenews.com.pk
unknown
wtvbamChina’s Z.ai says new model nears Anthropic’s Mythos 5 in cyber-defence tests
unknown
channelnewsasiaChina's Z.ai says new model nears Anthropic's Mythos 5 in cyber-defence tests - CNA