Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
China's Z.ai Says GLM-5.3 Beat Anthropic's Restricted Model on Bug-Hunting Test, But Still Lags on Actual Exploits

Since Z.ai unveiled GLM-5.3 on August 14, the headline claim has spread across tech coverage in the U.S., India, and China: a Chinese open-weight AI model scored higher than Anthropic's restricted Mythos 5 on a cybersecurity benchmark called CyberGym.
The number is real, as far as it goes. Z.ai reported GLM-5.3 hit 84.5% on CyberGym, a test that measures whether an AI can read source code, find security flaws, and confirm they're genuine. Anthropic's Mythos 5 scored 83.8%. OpenAI's GPT-5.6 Sol came in at 83.6%, according to Z.ai's own release page.
That's a genuine improvement over Z.ai's prior model. GLM-5.2 scored 77.2% on the same test back in June, according to techstartups.
But finding a vulnerability and actually weaponizing it are two different skills. On ExploitBench, which tests whether a model can build a working exploit from a known flaw, GLM-5.3 scored 54.4%. Mythos 5 scored 78%. GPT-5.6 Sol scored 76.5%. In a six-hour timed exploit-development test, GLM-5.3 completed 130 attack tasks. Mythos 5 completed 247, nearly double, according to both eWeek and the South China Morning Post.
GLM-5.3 is good at spotting problems. Anthropic's model is still far better at breaking things open.
Who's checking the math
Every performance number in this story traces back to one source: Z.ai's own release materials. Techstartups.com stated this plainly: "Z.ai's claims have not been independently verified." That's an important caveat that some coverage buried or skipped.
The model's actual weights weren't even available when Z.ai made the announcement. The company said it would delay the public release by roughly two weeks to run safety audits and security hardening, according to the Times of India and eWeek. That means outside researchers couldn't immediately test the claims themselves. Until they do, this is a vendor's self-reported scorecard, not a peer-reviewed result.
Z.ai also announced a "Security Disclosure Ledger" listing 2,436 vulnerabilities the model found across 269 open-source projects, including Linux kernel components and Apache software, with 1,097 rated medium-to-high severity, according to eWeek and SCMP. Some of those bugs reportedly date back roughly 40 years, according to SCWorld, citing The Register.
The commercial reality check
Strong benchmark numbers didn't help Z.ai's stock. Shares of the Hong Kong-listed company fell nearly 4% the day of the announcement, according to eWeek. Bloomberg Intelligence analyst Robert Lea was blunt about why: "This firm remains on a completely unsustainable commercial footing," he said, adding that "Rising agentic AI will drive Z.ai's inference costs and losses higher."
A model can post good benchmark numbers and still bleed cash. Building and running frontier-scale AI, especially open-weight models that give away the underlying technology, is expensive, and Z.ai's own market is signaling skepticism about whether the business model works long-term.
The safety framing is doing a lot of work
Z.ai is marketing its staggered rollout, publishing benchmark claims now, releasing weights later, and gating the most dangerous exploit capabilities behind a "trusted access" vetting program, as evidence of responsible AI development. Gabriel Wagner, an AI governance researcher at Concordia AI, told Reuters that "this is the first time a Chinese lab is publicly justifying a delayed open release of model weights with safety considerations," calling it a sign that "open-weight risk management practices in China are becoming more sophisticated."
That's a fair, on-the-record assessment. Anthropic built its entire Mythos program around the opposite instinct: keeping the most capable version locked to vetted partners under its "Project Glasswing" initiative rather than open-sourcing anything. Z.ai is trying to have it both ways, claiming both openness and caution, and it's Z.ai's own PR describing its own restraint.
What's actually unresolved
Nobody outside Z.ai has run these benchmarks yet. The 700-billion-parameter base model behind GLM-5.3 is the same one used in GLM-5.2, according to Z.ai, meaning the company says all of the reported gains came from additional reinforcement learning and post-training, not a bigger model. If that holds up under independent testing, it would matter more for the industry than the topline CyberGym score, since it would mean labs can extract major capability jumps without another multibillion-dollar pretraining run. Z.ai has not announced a firm date for the weight release beyond the roughly two-week window it cited on August 14, so that verification hasn't happened yet.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.