aiexpert
Home / News / Brief
Research · Aug 18, 2026, 04:06 PM · 4 sources

Z.ai releases GLM-5.3 with 84.5% CyberGym score; delays open weights for safety review

Z.ai released GLM-5.3 on August 14, 2026, a 743-billion-parameter model built on GLM-5.2's base with all gains from post-training alone. The model delivers 66.9% on DeepSWE coding (+20 points from predecessor) and a headline 84.5% on CyberGym vulnerability discovery, slightly edging Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%), per multiple sources.

Notably, Z.ai delayed open-weight release, departing from GLM-5.2's fast rollout pattern. The model found 2,436 vulnerabilities across 269 open-source projects— 1,097 rated critical or high severity— surfaced through coordinated disclosure. Z.ai attributes the surge to emergent exploit-chain reasoning the company says it did not explicitly train for, citing the need for safety evaluation before public weights.

The move carries policy significance: it is the first time a Chinese frontier lab cited emergent capability concerns (rather than export rules or platform policy) for withholding weights. For architects evaluating open models, GLM-5.3 marks the inflection where vulnerability discovery (defensive) scales faster than open-source deployment guardrails. The two-week hold and future weight release, expected end-August, will be when independent verification begins on the company's benchmark claims.

Sources

Everything this brief rests on
  1. 01 Primary source unite.ai
  2. 02 Z.ai Launches GLM-5.3 With Frontier Coding and Cyber Capability unite.ai “Z.ai released GLM-5.3 on August 14, 2026... keeps the same base model as GLM-5.2 and derives every capability gain from scaled-up post-training... GLM-5.3 is the strongest open-weights system it has measured”
  3. 03 AI Weekly aiweekly.co “GLM-5.3 scored 84.5% on CyberGym... ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%... Z.ai says the model surfaced 2,436 vulnerabilities across 269 open-source projects, 1,097 rated critical or high, and delayed weights about two weeks”
  4. 04 explainx.ai Blog explainx.ai “GLM-5.3 did not follow that script... Right now... 'An initial group of partners is now offering GLM-5.3-powered services through our official service, with its safeguards and usage policies in place'”