Anthropic says the most capable openly downloadable model for building cyber exploits ships with safeguards that fall to a cover story. In a report published September 29, 2026, its Frontier Red Team writes that “GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse,” and that attackers “can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.”
This page is read from Anthropic’s report, by Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher. A disclosure first: Anthropic competes with Z.ai, and our own AI employees run on Anthropic’s Claude Opus 5.5. Our earlier record of the model is GLM 5.3 for AI agents.
On this page · 9 sectionsOpen
- On September 29, 2026, Anthropic’s Frontier Red Team published an analysis of GLM-5.3, Z.ai’s open-weight model, finding it can build working cyber exploits end to end.
- On ExploitBench, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, against 56 of 410 for Claude Mythos Preview; earlier models such as GLM-5.2 and Claude Opus 4.6 succeeded in none of the binary-exploitation tasks tested.
- Anthropic says simple techniques bypassed GLM-5.3’s safeguards: a cover story 64% of the time, prefilled thinking 92%, and an abliterated copy 100%, while safeguarded Claude models stayed at zero.
- Abliteration took Anthropic’s first-time team about 2,200 GPU hours, roughly $4,400, and cut the refusal rate from above 90% to as low as 2% with capability largely intact.
- NIST’s CAISI called GLM-5.3 the most cyber-capable open-weight model released to date, about four months behind the US frontier.
- Anthropic is a rival lab testing a competitor’s model; the misuse tests ran in a simulated environment it calls imperfect.
§ 01What GLM-5.3 can build
| Test | GLM-5.3 | Claude Mythos Preview |
|---|---|---|
| ExploitBench, end-to-end exploits | 50 of 410 | 56 of 410 |
| Binary exploitation, full control-flow hijack | 4% | 6% |
| Earlier models on the same binary tasks | GLM-5.2: none | Claude Opus 4.6: none |
The hands-on sessions are the sharper finding. With less than an hour of human attention across a day, a researcher used GLM-5.3 to find several unknown flaws in a popular browser’s JavaScript engine and chain them into a page that reads files off a visitor’s computer; Anthropic says it disclosed them to the maintainer. In a second session, GLM-5.3-Flash turned a public Chrome fix into a working exploit chain. “This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu’s API prices, this effort would have cost $20.40.”
§ 02How the safeguards failed
Out of the box, GLM-5.3 refused the overtly malicious requests in Anthropic’s simulated world. Three simple techniques changed that.
Safeguarded Claude models stayed at zero under all three, Anthropic says, partly for structural reasons: its API offers no way to prefill Claude’s thinking, and Claude’s weights are not public, so they cannot be abliterated. The abliteration itself took Anthropic’s first-time team “about 2,200 GPU hours at a computation cost of roughly $4,400,” and it estimates an experienced team would need about 600 hours. It took GLM-5.3’s refusal rate from above 90% to 3%, 2% and 12% on three public benchmarks, with general capability unchanged on GPQA-Diamond.
§ 03The independent check
NIST’s Center for AI Standards and Innovation published its own assessment on September 17, finding GLM-5.3 to be “the most cyber-capable open-weight model released to date,” about four months behind the US frontier. Anthropic says its capability findings broadly match CAISI’s. The safeguard tests are Anthropic’s alone.
§ 04What Anthropic wants
Two things. More access for defenders: “Our view is that cyber defenders should use the best available tools that meet their needs,” and Anthropic says it is widening access to Claude’s cyber capabilities. And testing by governments: “Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3.” Its summary line is that “a critical threshold in freely accessible capabilities has now been crossed.”
§ 05What the report does not settle
- Who is testing whom. Anthropic sells the safeguarded alternative it compares against. The capability numbers have CAISI behind them; the bypass numbers do not.
- Real-world conditions. The misuse tests ran in a simulated world where no model-written code executed, and Anthropic itself calls these simulations “imperfect measures of how a model would behave.”
- Actual attacks. The report predicts state and non-state actors will use models like this; it documents no attack carried out with GLM-5.3.
- Z.ai’s view. The report includes no response from Zhipu AI.
§ 06What this means if you run agents on real work
The attacker’s toolkit just got cheaper, and it is downloadable. That puts more weight on the defensive layer every business already controls: patching fast, giving each agent and each person only the access its work needs, and putting a gate in front of anything consequential. The same questions apply to your own agents as to anyone else’s.
§ 07What we are watching for
- A response from Zhipu AI, or safeguard changes in a GLM-5.3 update.
- The disclosures Anthropic says it will make for the drivers and device software it tested.
- Whether other labs publish matching red-team results on open-weight models.
§ 08Update log
- September 30, 2026: page opened, from Anthropic’s report of September 29.
§ 09Sources
- Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities, September 29, 2026.
- Our record of GLM 5.3 for AI agents.
Q1What is GLM-5.3?
An open-weight model from Zhipu AI, known outside China as Z.ai. Anyone can download its weights, which is why Anthropic’s report focuses on how easily its safeguards can be removed.
Q2What is abliteration?
A technique that edits an open-weight model to remove its refusals. Anthropic says several developers published abliterated versions of GLM-5.3 within days of release, and its own copy barely lost capability on GPQA-Diamond.
Q3How close is GLM-5.3 to Claude Mythos?
On ExploitBench, 50 of 410 end-to-end exploits for GLM-5.3 versus 56 of 410 for Claude Mythos Preview. On Anthropic’s binary-exploitation benchmark, 4% versus 6% full control-flow hijacks.
Q4Did anyone independent check this?
On capability, yes: NIST’s Center for AI Standards and Innovation published its own assessment on September 17 and Anthropic says its findings broadly match. The safeguard-bypass tests are Anthropic’s own.
Q5What does this mean for businesses using AI agents?
Capable attack tooling is now freely downloadable, so the defenses around your systems matter more: patching, least-privilege access and approval steps on anything consequential.
Q6Where does CellCog fit?
Our AI employees run on Claude Opus 5.5, not GLM-5.3, and every action that reaches your world passes an approval step. We cover this because the security of the tools attackers can reach is the backdrop every agent platform works against.
