GLM-5.3: The Open-Weights Coding Model With a Cyber Edge

Z.ai released GLM-5.3, the most capable open-weights coding model yet, with cyber skills that beat Mythos 5 on CyberGym. Weights land in two weeks.

GLM-5.3: The Open-Weights Coding Model With a Cyber Edge

Z.ai just released GLM-5.3, and it’s the most capable open-weights coding model the company has shipped yet. The model shares the same base as GLM-5.2, so every gain comes from post-training. Same base, same architecture, same context window, and a noticeably better model on paper.

It’s the coding model that could reset expectations for what post-training alone can do.

What makes this coding model different

Z.ai says the gains came from training, not from a new brain. The company says GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench.

On public benchmarks the gains are dramatic: Terminal-Bench 3.0 jumps from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents’ Last Exam from 23.8 to 28.5. Those are open-source SOTA numbers on Terminal Bench 3.0 and Agents’ Last Exam, which is exactly the territory where the best AI coding agents I’ve ranked this year live.

The efficiency story is the sleeper. At Max effort, GLM-5.3 reaches 34.5% at roughly 75K output tokens per task, compared with 23.4% at 96K for GLM-5.2. Better results with fewer tokens is the direction every serious coding model needs to head, and Z.ai is pointing that way.

That matters more than the leaderboard, because tokens are the bill you actually pay when you run agents all day. A coding model that does more with less is the kind of improvement you feel in the wallet, not just in the charts. The token math is the part most launch posts bury, and Z.ai put it on the front page.

The cyber surprise

Here’s where it gets uncomfortable, in a good way. As Z.ai scaled post-training it introduced vulnerability discovery data and environments into the training mix, and the capability grew faster than expected. Cyber capability emerged as a side effect of scaling post-training, and the company is honest that it surprised them.

GLM-5.3 scores 84.5% on CyberGym, up from GLM-5.2’s 77.2%, ahead of Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

The further up the exploitation chain you go, the bigger the jump. On ExploitBench, GLM-5.3 reaches 54.4%, more than doubling GLM-5.2’s 24.4%, though Mythos 5 still leads at 78.0%.

On ExploitGym, it completes 105 tasks within two hours and 130 within six, compared with 29 and 39 for GLM-5.2. Mythos 5 remains well ahead at 181 and 247 tasks, so the closed frontier still holds the top rung.

The pattern tells you something useful about where open weights actually stand. Finding flaws is one thing, and turning them into working attacks is another, much harder thing. That’s the honest gap in the whole release, and it’s a big one.

That capability isn’t just benchmark theater. Working with security teams in China, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues.

The findings span kernels, browsers, open-source infrastructure, and network protocols, and the oldest dates back roughly 40 years. The oldest flaws predate the modern internet, which is older than most of the software industry can claim.

Z.ai now runs a Security Disclosure Ledger that tracks the findings as they move through disclosure, with 53 already public and 2,383 under embargo. A model vendor publishing a running ledger of real vulnerabilities it found is rare, and it changes how you read the benchmark claims.

The ledger also gives you a way to check the model’s work instead of taking the numbers on faith.

The open-weights catch

Here’s the part I want you to hold onto. Z.ai will release the weights in two weeks after launch, once safety evaluation and hardening are complete. The model is live now through the GLM Coding Plan and ZCode, but the open-weights part of the story is a promise, not a download.

The most sensitive cyber functions will also sit behind a trusted access program for verified users, and Z.ai is starting an Open Source Shield initiative to audit open-source projects and hand model access to defensive teams.

Anthropic keeps its cyber-capable Mythos model locked behind vetted access, and according to Gabriel Wagner, an AI governance researcher at Concordia AI, this is the first time a Chinese lab has publicly justified a delayed open release of model weights with safety considerations.

The comparison is doing real work. GLM-5.3 is a general-purpose coding model that picked up cyber skills through post-training, not a purpose-built security system, which makes the safeguards conversation more interesting, not less.

It’s also the first time a Chinese lab has publicly held back open weights for safety reasons, and that shift is worth watching.

Watch what the trusted access program actually gates, because that’s where the real policy lives.

Should you care?

Yes, if you build with agents. GLM-5.3 is the open-weights coding model to watch this month, and the benchmark pattern matches what I’ve seen in the undetected agent hacks and CyberGym coverage: capability is growing fastest exactly where the open models are furthest behind.

Just remember the numbers are Z.ai’s own, and nothing is independently verified until the weights land and the community reruns the tests. That two-week window is the honest part of the launch.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *