Inside Zhipu’s New AI Model: Where It Beats American Rivals and Where It Trails.
Zhipu's GLM-5.3 Claims a Cybersecurity Win — But Its Own Numbers Tell a Split Story The headline benchmark edged out Anthropic and OpenAI. The two benchmarks nobody covered show the opposite result.
84.5%: CyberGym Score, Leads the Field
54.4%: ExploitBench, Trails by 24 Points
2,436: Vulnerabilities Found in Testing
1: The Headline Versus the Fine Print:
Zhipu, also known as Z.ai, released GLM-5.3 on August 14 as a coding-focused model, publishing a technical note detailing how it stacks up against rival systems. The claim that spread fastest concerned security: GLM-5.3 scored 84.5% on a vulnerability-discovery benchmark called CyberGym, edging out Anthropic's Mythos 5 at 83.8% and OpenAI's GPT-5.6 Sol at 83.6%. Coverage ran with the framing that a Chinese model had overtaken its American rivals at finding software bugs.
Zhipu's own release is more careful than the headlines it generated. The CyberGym result is real, but it's also the narrowest of three cybersecurity benchmarks the company published — and the note is upfront that the other two point the other way. Buried in the same release is a line that sums up the pattern better than any headline did: GLM-5.3's security capability, Zhipu writes, is growing fastest exactly where the model is furthest behind.
2: Three Benchmarks, Three Different Pictures:
CyberGym tests whether a model can read source code and confirm a genuine vulnerability. That's the test GLM-5.3 narrowly won, by seven tenths of a percentage point. ExploitBench asks something tougher — reasoning through how a real vulnerability would actually be exploited — and here GLM-5.3 scores 54.4%, more than double its predecessor's 24.4%, but still well behind Mythos 5's 78.0% and GPT-5.6 Sol's 76.5%.
The third test widens the gap further. ExploitGym measures how many exploitation tasks a model can finish within a fixed time window. GLM-5.3 completes 105 tasks in two hours and 130 in six; Mythos 5 completes 181 and 247 in the same windows. Finding a flaw and building a working exploit from it are different skills, and Zhipu's own data shows the model falling further behind the closer a test gets to the second one.
3: A Moving Comparison and a Common Harness:
Part of the confusion comes from which Anthropic model gets used where. GLM-5.3's main benchmark table compares it to Opus 4.8, its performance charts use Fable 5, and its cybersecurity section uses Mythos 5 — three different comparisons in three different places, easy to blur into one. On coding generally, the result is mixed: GLM-5.3 beats Opus 4.8 on some tasks and loses on others, and Zhipu states directly that its model still trails Claude Fable 5 on its own internal coding benchmark.
The testing method is also worth a note. Zhipu ran GLM-5.3 and its rivals through Claude Code 2.1.207, Anthropic's own coding agent, as the common harness for CyberGym, ExploitGym, ExploitBench, and other tasks. That's a reasonable way to keep a comparison fair, and Zhipu documents its settings — but it's a reminder that a Chinese open-weights model's frontier claims are still being measured on American tooling.
Cybersecurity capability is growing fastest exactly where we are furthest behind. — Zhipu AI, GLM-5.3 technical release note
4: What the Vulnerability Count Doesn't Say:
Beyond the benchmarks, Zhipu says it worked with security teams to run its models against real-world codebases, surfacing 2,436 vulnerabilities across 269 open-source projects — 107 critical, 990 high, 1,286 medium, and 53 low. The average flaw had gone unnoticed for 26.6 years, with the oldest dating back to 1981. Of the total, only 53 have been publicly disclosed; the remaining 2,383 stay under embargo.
One inconsistency is worth flagging: Zhipu's own summary panel labels 1,097 findings as "critical and high," matching its severity table, while the body text of the same release calls that same figure "medium-to-high" — a discrepancy several outlets have carried forward without noticing.
The release also doesn't say how many of the 2,436 findings were previously unknown or independently reproduced, which are the two numbers that would turn a volume claim into a proven capability claim. The count itself is post-review and deduplicated, not raw model output.
5: The Numbers That Matter More Than the Benchmark:
Two details in the release carry further than the CyberGym margin. The first is efficiency: GLM-5.3 reaches 31.4% on Zhipu's internal coding benchmark using roughly 50,000 output tokens per task, versus Opus 4.8's 29.5% at 120,000 tokens — noticeably better results for under half the cost, which matters most to teams without frontier-scale budgets.
The second is distribution. Zhipu says it will publish GLM-5.3's weights once safety evaluation and hardening are complete — which hasn't happened yet, making the open-weights promise a commitment rather than a fact today. If it holds, a model with documented vulnerability-discovery ability becomes downloadable and runnable locally by any team, including in markets that will never get access to an export-controlled American model.
The Real Lesson: One Headline Number Is Never the Whole Story.
GLM-5.3's release shows how easily a single strong benchmark can outrun the fuller picture — a company can lead on one measure and trail by double digits on the next, and the difference only shows up if someone reads past the headline. Enterprises evaluating AI tools face the same trap every day: a flashy demo or a favorable number doesn't tell you whether a system holds up in production.
Otherworlds AI's Agent+ platform is built for that reality — production-tested AI agents on Google Opal automated workflows, starting at $297/month, with custom enterprise AI builds for teams that need a deeper fit. No benchmark-chasing, just AI that performs where it counts.
Explore Agent+ at otherworldsai.com







