What happened
- On September 29, Anthropic published an analysis of GLM-5.3, the model from China’s Zhipu AI (Z.ai outside China). It concludes that the model builds exploits, meaning code that takes advantage of a security flaw, end to end and autonomously, with a capability close to that of Claude Mythos Preview, the model Anthropic presented five months earlier as the first to do so.
- The difference is not in capability but in the brakes. Anthropic says GLM-5.3 shipped without meaningful safeguards and that its simulated tests get past them between 64% and 100% of the time with simple techniques. Against Claude models with safeguards, the same techniques did not work.
- The model’s weights are published on Hugging Face and anyone can download them. Z.ai launched it on August 14 and, according to the coverage we reviewed, held back the weights for about two weeks for a “safety evaluation and hardening,” without publishing a system card or a safety report.
- On September 17, the Center for AI Standards and Innovation (CAISI) at NIST, the U.S. standards agency, had reached a similar conclusion by another route: GLM-5.3 is the most cyber-capable open-weight model released to date, though it trails the U.S. frontier by about four months.
What Anthropic measured
Capability. On ExploitBench, which measures whether a model can turn known flaws in Chrome’s V8 engine into a complete exploit, GLM-5.3 managed it in 50 of 410 attempts (12%) and Mythos Preview in 56 (14%). On an internal test of exploiting flaws in open-source projects, with full credit only if control of the program’s flow is taken, GLM-5.3 succeeded in 4% of 100 attempts and Mythos Preview in 6%. Earlier models, such as Claude Opus 4.6 and GLM-5.2, succeeded in none, and on ExploitBench Kimi K3 (from Moonshot AI) and DeepSeek V4.1-Flash came in at zero or close to it.
Real work. A researcher used GLM-5.3 in an isolated Linux environment and, in a single day with little human attention, found and chained together previously unknown flaws in the JavaScript engine of a popular browser. The result was a web page that, when visited, reads arbitrary files from the visitor’s computer, including SSH private keys. The model also found exploitable flaws in wireless and graphics drivers and in network device software. The browser ones have already been reported to their maintainer; the rest are under review before disclosure. Another researcher gave GLM-5.3-Flash, the smaller version, the public details of a recent Chrome flaw and of another known flaw; with no significant direction, the model chained them and produced a reliable exploit for an ARM64 processor that gets around pointer-authentication protection. It took 20 minutes of human attention and 8 hours of the model’s work, which at Z.ai’s API prices would have cost US$ 20.40.
Brakes. Given a direct malicious order, GLM-5.3 always refused (0%). Given a false cover story, for example telling it that it is an agent in a red-team exercise, it took part 64% of the time. With its reasoning pre-filled so that it appears to have already decided to go ahead, 92%. And with an “abliterated” version, 100%.
Abliteration is a well-known technique that erases from the model the internal direction that produces refusals. Anthropic replicated it: it cost about 2,200 GPU hours, close to US$ 4,400, with no prior experience; a skilled team would need about 600 hours, US$ 1,200. The refusal rate fell from more than 90% to about 3% on JailbreakBench, and capability barely changed. According to Anthropic, several developers published abliterated versions of GLM-5.3 within days of the launch.
Why it matters
- A lead of months is over, or nearly. For months, the ability to build exploits autonomously sat in the hands of a few labs and behind restricted-access programs. Now there is an open version, from another country, that nearly matches the Mythos of that time and trails the current frontier by about four months.
- Control no longer lives in the model. With Claude, the brake is in the channel: an API that doesn’t let anyone pre-fill the reasoning, and weights nobody can abliterate. An open model has no such channel, and every safeguard built into the weights is a suggestion that can be undone with hundreds of GPU hours.
- No actor has a clear lever to stop it today. The weights are already published and copied. Anthropic asks that developers of open models around the world safeguard them adequately, and that governments test models of sufficient capability. That is a wish and a recommendation, not a mechanism anyone can enforce.
Two measurements that don’t contradict each other
At first glance, the two documents don’t add up. Anthropic says GLM-5.3 nearly reaches Mythos Preview; NIST says it trails the U.S. frontier by four months and shows large gaps: 40.4% against 90.2% at finding flaws in browser engines (SEC-Bench Pro), 9.4% against 44.4% at developing exploits for real flaws (ExploitGym) and 7.7% against 23.2% on its private test with open-source projects.
Both things are true because they compare against different references. Mythos Preview is five months old; the current frontier, which includes Mythos 5 and 5.1, has moved on since then. The reading that reconciles both documents is this: the open model is chasing the closed ones with a delay of about four months, according to NIST. It is a window, not a permanent gap. NIST also tested the U.S. models, where applicable, with their cybersecurity safeguards switched off and reports no equivalent conditions for GLM-5.3, so the pure-capability comparison measures something different from what an attacker can do today with each one.
The race changes shape
From wall to window. The U.S. strategy, as Anthropic’s text shows it, rests on three pieces: closed weights, verified access for defenders and a window of advantage so that defenders can patch first. With the Glasswing program, Anthropic says, those defenders found more than 10,000 vulnerabilities in critical software before an equivalent model without controls existed. GLM-5.3 shows how long that window can last in practice: about four months, according to NIST.
Two governance models on the same track. One controls who gets access and can therefore give preference to defenders. The other publishes the weights and hands control to whoever downloads them. Neither solves the problem on its own: the first depends on the vetting process being good and on no one else reaching the same level; the second depends on good faith prevailing. The September 8 alert from the NSA, FBI and CISA, which already named Z.AI among the companies that allegedly distilled U.S. models, shows that Washington already frames competition with Chinese labs as a national security matter. GLM-5.3 shifts the concern from how a model is trained to who uses it and for what.
Hardware is no longer the only bottleneck. The model card on Hugging Face says the model can be deployed on platforms with Ascend NPUs, Huawei’s chips, using the vLLM-Ascend, xLLM and SGLang frameworks. It is a technical line, not a political statement, but it suggests that a Chinese lab can serve a model of this class without depending entirely on U.S. hardware. That is our inference from that mention, not a fact confirmed by Anthropic’s text.
And the repository is American. The weights are on Hugging Face, a platform that Nvidia confirmed it will buy for US$ 12.93 billion. Anthropic does not propose pulling the model: it asks for more government testing and safeguards from developers. NIST only publishes its assessment. How much decision-making power a platform has over what it hosts is a question the race has just put on the table.
On the wrong side of the gap
Anthropic says verified defenders can already use models such as Mythos 5.1 through its trusted-access programs, and proposes extending access to more entities. In the text we reviewed we did not find the criteria for being verified or a list of countries. That silence is the political point. A Chilean bank, a power company or the incident-response team of a public agency compete against attackers who today can download the same class of tool without asking permission. Whether they can use the closed version depends on decisions made in San Francisco.
For countries that don’t build these models, the race has a practical consequence: they are split between two dependencies. Either they depend on a foreign provider that decides who counts as a defender, or they are exposed to an open model that any attacker uses without conditions. Chile has a legal framework and an agency for cybersecurity (the Cybersecurity Framework Law and the ANCI) and a personal data law with obligations that will apply to whoever suffers a breach. Nothing published indicates that this framework includes guaranteed access to frontier defensive tools.
What this means for security careers
This part is our reading, not the source’s, and it is worth saying so.
- What gets cheaper is the step from “finding” to “having an exploit.” US$ 20.40 and 20 minutes of human attention for an exploit against an already disclosed flaw change the economics of second-hand attacks: the ones that exploit a patch the victim hasn’t applied yet. What used to demand a specialized researcher’s time now demands an afternoon and a credit card.
- What stays expensive is deciding what matters. In both of Anthropic’s cases, a researcher set the task, checked the result and decided what to do with it. The capability scales; the judgment to prioritize, validate and disclose is still human.
- The pressure lands first on whoever patches. If the time between releasing a patch and seeing a working exploit shrinks to hours, the bottleneck moves to the teams that have to apply updates across dozens of systems. The most valuable profile stops being the person who “knows how to attack” and becomes the one who can reduce attack surface, automate patching and respond to an incident with few hands.
- For those starting out, the entry-level rung moves. The tasks of a junior penetration-testing analyst, such as reproducing known flaws or writing proofs of concept, are exactly what the model does for a few dollars. The flip side is that the tool also exists on the defensive side for whoever learns to use it.
What the source doesn’t say
- Anthropic is not a neutral observer. It sells the closed alternative, proposes expanding access to its own models and asks for more government testing. It can be right and still have an interest in the conclusion.
- The safeguard tests are simulated. According to the post itself, no model-generated code is ever executed and a second model approximates the result of each command. They measure willingness to take part, not real attacks.
- Z.ai’s numbers are its own. The 84.5% on CyberGym, above Mythos 5 (83.8%), and the 2,436 vulnerabilities in 269 open-source projects come from the company itself, with no independent reruns to date.
- There is no response from Z.ai in what we reviewed, and Anthropic’s text doesn’t say whether it was notified before publication.
The number
What’s next
- Anthropic asks governments to test models of sufficient capability and their successors. No dates or mechanism announced.
- GLM-5.3’s successors are the next decision point. If the lag of about four months holds, the next open generation would reach what is today’s frontier around early 2027. That is an extrapolation from two measurements, not anyone’s forecast.
- Anthropic proposes extending access to frontier models to more defending entities, with no figures on how many organizations or from which countries.
Bottom line
Five months ago, the ability to build an exploit without human help was an industrial secret with a waiting list. Today it is a file you download, and a third party takes the brakes off it with a few thousand dollars of compute. The race is no longer just about who gets there first, but about how long the advantage lasts for whoever did. The answer, for now, is measured in months.
Sources
- GLM-5.3 and the spread of advanced cyber capabilities, Anthropic, September 29, 2026.
- CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities, NIST, September 17, 2026.
- zai-org/GLM-5.3, model card on Hugging Face.
- Z.ai ships GLM-5.3, holds open weights for cyber safety review, AI Weekly, August 15, 2026.
- Anthropic warns a free Chinese AI model can already build working hacks, Startup Fortune, September 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select, verify and correct each note.




