AERIOXFLUX
Frontier Labs
Frontier Labs · chinese labs

GLM-5.3 Found 2,436 Vulnerabilities. Then Z.ai Held the Weights.

A Chinese lab shipped a 743B open-weights coding model that tops a cybersecurity benchmark, surfaced a flaw in Cursor, and delayed its own release by two weeks for safety review — the first time an open-weights lab has voluntarily done that on capability grounds.

Flux Desk·2026-08-16·5 min read

Within a day of GLM-5.3 going live, the model had found a significant vulnerability in Cursor, the AI code editor — and Z.ai was publishing benchmark numbers that put its 743-billion-parameter open-weights release at the top of a cybersecurity evaluation, ahead of models from Anthropic and OpenAI.

The coding benchmarks got the coverage. The security numbers are the ones that change what an open-weights release means.

On CyberGym, a vulnerability-discovery benchmark, Z.ai reports 84.5% — ahead of Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. The lab says the model surfaced 2,436 vulnerabilities across 269 open-source projects, with 1,097 rated critical or high.

And then it did something no open-weights lab has done before on these grounds: it held the weights for roughly two weeks to run a safety evaluation before release.

The capability is a side effect, and Z.ai says so

GLM-5.3 was not built as a security model. It is a coding model, and the security performance is what Z.ai describes as emerging from post-training reinforcement learning in security-focused environments — the same general recipe that produces better software engineering also produces better vulnerability discovery, because the two tasks are closer than the industry's org charts suggest.

Finding a bug and finding an exploitable bug are the same reasoning process run to different depths. A model that can hold a large codebase in context, trace data flow across module boundaries, and reason about what a function assumes about its inputs is doing offensive security whether or not anyone labeled the training data that way.

This is the specific thing that safety researchers have warned about for two years: capability in this domain is not a separate axis you can decline to train. It arrives attached to the thing everyone wants.

The two-week delay is the story

Open-weights releases have historically been governed by a simple norm — ship the weights with the paper, or shortly after, because the entire value proposition is that nobody gates access.

Z.ai broke that norm deliberately. It shipped the API and the benchmarks first, then held the downloadable weights for a safety review before releasing them.

That is a Chinese lab voluntarily adopting a staged-release practice that Western labs spent years arguing about and mostly abandoned for open models. It is worth being precise about what it is and is not. It is not a refusal to release — the weights are coming. It is not third-party audited. It is not a binding commitment to do the same next time.

But it is an open-weights lab publicly conceding that a capability threshold exists past which "just ship it" stops being the obvious answer. Whatever the motivation — genuine caution, regulatory anticipation, or the reputational math of being the lab whose model gets credited in an incident report — the precedent is now set by a lab that had no obligation to set it.

The Cursor finding is the demonstration that matters

Benchmarks are contested. A model finding a real vulnerability in a widely deployed developer tool is not.

GLM-5.3 identified a serious flaw in Cursor — a product with enormous install base among exactly the population that would also be running GLM-5.3. That is a useful demonstration and an uncomfortable one, in the same gesture: the tool discovers flaws in the tools, and the population with access to the discovery is identical to the population with access to the codebases.

2,436 findings across 269 projects is a volume number, and volume numbers in vulnerability discovery deserve scrutiny — static analyzers have generated large finding counts with poor precision for decades. The 1,097 critical-or-high figure is vendor-rated, not independently triaged. Assume some fraction are duplicates, some are false positives, and some are real-but-unreachable.

Even discounted heavily, the remainder is a lot of live bugs in open-source code, found by one model, in one pass.

The numbers are vendor-published, and that limits what they prove

Every figure here originates with Z.ai or is cited from Z.ai. There is no single independent run under a common harness, a common compute budget, and a common context setting.

That caveat is not a formality. Benchmark results in this category are extremely sensitive to scaffolding — how many attempts the model gets, how much of the codebase it sees, whether it can execute code, how long it runs. A 0.7-point margin over Mythos 5 is well inside the range that harness differences produce. The honest reading of the CyberGym result is "competitive with the frontier," not "beat the frontier."

The 54.4% figure Z.ai reports on exploit execution — reasoning through and actually running an exploit rather than just spotting the flaw — is the more interesting number, and it is lower for a reason. Identifying a vulnerability is pattern recognition over code. Weaponizing one requires a working model of the runtime, the memory layout, and the mitigations in place. The gap between the two scores is the gap between a research capability and an operational one.

The read

Three things follow.

Frontier vulnerability-discovery capability is now open-weights. Not "will be" — is. The gating question for the last two years was whether the models best at finding exploitable bugs would stay behind APIs with usage monitoring and abuse review. That question is answered. Downloadable, runnable, unmonitored, in the same weight class as the closed frontier.

The defensive case is real and asymmetric in an underappreciated way. The same model that helps an attacker helps a maintainer, and maintainers of open-source projects have historically had no security budget at all. A free model that finds critical bugs in a dependency nobody was paying to audit is a genuine improvement over the status quo, which was that nobody looked. The asymmetry cuts both ways — but the defensive side starts from a much lower base.

The release norm just moved, and it moved from an unexpected direction. The Western policy conversation has assumed Chinese labs would be the ones racing to release without gates. Z.ai just delayed weights for a safety review of its own accord. The follow-up question is not whether that was sincere. It is whether it becomes the practice or stays the exception — and the answer arrives with the next model, not this one.

#z-ai#glm-5-3#open-weights#cybergym#vulnerability-discovery

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.