A closed vehicle checkpoint barrier in the middle of a plain, with dirt footpaths going around it on both sides, under an orange sun disc

Everyone called for slowing the pace and kept going:

An open Chinese model nearly matches Mythos from five months ago at building exploits, and its guardrails come off for US$ 4,400


Anthropic showed that GLM-5.3 builds end-to-end exploits and that its safeguards give way between 64% and 100% of the time. The U.S. lead has gone from a wall to a window: about four months.

September 29, 2026 · Translated from the Spanish original

What happened

What Anthropic measured

Capability. On ExploitBench, which measures whether a model can turn known flaws in Chrome’s V8 engine into a complete exploit, GLM-5.3 managed it in 50 of 410 attempts (12%) and Mythos Preview in 56 (14%). On an internal test of exploiting flaws in open-source projects, with full credit only if control of the program’s flow is taken, GLM-5.3 succeeded in 4% of 100 attempts and Mythos Preview in 6%. Earlier models, such as Claude Opus 4.6 and GLM-5.2, succeeded in none, and on ExploitBench Kimi K3 (from Moonshot AI) and DeepSeek V4.1-Flash came in at zero or close to it.

Real work. A researcher used GLM-5.3 in an isolated Linux environment and, in a single day with little human attention, found and chained together previously unknown flaws in the JavaScript engine of a popular browser. The result was a web page that, when visited, reads arbitrary files from the visitor’s computer, including SSH private keys. The model also found exploitable flaws in wireless and graphics drivers and in network device software. The browser ones have already been reported to their maintainer; the rest are under review before disclosure. Another researcher gave GLM-5.3-Flash, the smaller version, the public details of a recent Chrome flaw and of another known flaw; with no significant direction, the model chained them and produced a reliable exploit for an ARM64 processor that gets around pointer-authentication protection. It took 20 minutes of human attention and 8 hours of the model’s work, which at Z.ai’s API prices would have cost US$ 20.40.

Brakes. Given a direct malicious order, GLM-5.3 always refused (0%). Given a false cover story, for example telling it that it is an agent in a red-team exercise, it took part 64% of the time. With its reasoning pre-filled so that it appears to have already decided to go ahead, 92%. And with an “abliterated” version, 100%.

Abliteration is a well-known technique that erases from the model the internal direction that produces refusals. Anthropic replicated it: it cost about 2,200 GPU hours, close to US$ 4,400, with no prior experience; a skilled team would need about 600 hours, US$ 1,200. The refusal rate fell from more than 90% to about 3% on JailbreakBench, and capability barely changed. According to Anthropic, several developers published abliterated versions of GLM-5.3 within days of the launch.

Why it matters

Two measurements that don’t contradict each other

At first glance, the two documents don’t add up. Anthropic says GLM-5.3 nearly reaches Mythos Preview; NIST says it trails the U.S. frontier by four months and shows large gaps: 40.4% against 90.2% at finding flaws in browser engines (SEC-Bench Pro), 9.4% against 44.4% at developing exploits for real flaws (ExploitGym) and 7.7% against 23.2% on its private test with open-source projects.

Both things are true because they compare against different references. Mythos Preview is five months old; the current frontier, which includes Mythos 5 and 5.1, has moved on since then. The reading that reconciles both documents is this: the open model is chasing the closed ones with a delay of about four months, according to NIST. It is a window, not a permanent gap. NIST also tested the U.S. models, where applicable, with their cybersecurity safeguards switched off and reports no equivalent conditions for GLM-5.3, so the pure-capability comparison measures something different from what an attacker can do today with each one.

The race changes shape

From wall to window. The U.S. strategy, as Anthropic’s text shows it, rests on three pieces: closed weights, verified access for defenders and a window of advantage so that defenders can patch first. With the Glasswing program, Anthropic says, those defenders found more than 10,000 vulnerabilities in critical software before an equivalent model without controls existed. GLM-5.3 shows how long that window can last in practice: about four months, according to NIST.

Two governance models on the same track. One controls who gets access and can therefore give preference to defenders. The other publishes the weights and hands control to whoever downloads them. Neither solves the problem on its own: the first depends on the vetting process being good and on no one else reaching the same level; the second depends on good faith prevailing. The September 8 alert from the NSA, FBI and CISA, which already named Z.AI among the companies that allegedly distilled U.S. models, shows that Washington already frames competition with Chinese labs as a national security matter. GLM-5.3 shifts the concern from how a model is trained to who uses it and for what.

Hardware is no longer the only bottleneck. The model card on Hugging Face says the model can be deployed on platforms with Ascend NPUs, Huawei’s chips, using the vLLM-Ascend, xLLM and SGLang frameworks. It is a technical line, not a political statement, but it suggests that a Chinese lab can serve a model of this class without depending entirely on U.S. hardware. That is our inference from that mention, not a fact confirmed by Anthropic’s text.

And the repository is American. The weights are on Hugging Face, a platform that Nvidia confirmed it will buy for US$ 12.93 billion. Anthropic does not propose pulling the model: it asks for more government testing and safeguards from developers. NIST only publishes its assessment. How much decision-making power a platform has over what it hosts is a question the race has just put on the table.

On the wrong side of the gap

Anthropic says verified defenders can already use models such as Mythos 5.1 through its trusted-access programs, and proposes extending access to more entities. In the text we reviewed we did not find the criteria for being verified or a list of countries. That silence is the political point. A Chilean bank, a power company or the incident-response team of a public agency compete against attackers who today can download the same class of tool without asking permission. Whether they can use the closed version depends on decisions made in San Francisco.

For countries that don’t build these models, the race has a practical consequence: they are split between two dependencies. Either they depend on a foreign provider that decides who counts as a defender, or they are exposed to an open model that any attacker uses without conditions. Chile has a legal framework and an agency for cybersecurity (the Cybersecurity Framework Law and the ANCI) and a personal data law with obligations that will apply to whoever suffers a breach. Nothing published indicates that this framework includes guaranteed access to frontier defensive tools.

What this means for security careers

This part is our reading, not the source’s, and it is worth saying so.

What the source doesn’t say

The number

US$ 20.40.It is what it would have cost, at Z.ai's API prices, to build a working exploit for a recent Chrome flaw with GLM-5.3-Flash: 8 hours of the model's work and 20 minutes of human attention.Source: Anthropic

What’s next

Bottom line

Five months ago, the ability to build an exploit without human help was an industrial secret with a waiting list. Today it is a file you download, and a third party takes the brakes off it with a few thousand dollars of compute. The race is no longer just about who gets there first, but about how long the advantage lasts for whoever did. The answer, for now, is measured in months.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select, verify and correct each note.

Related notes

← All notes