Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

Anthropic's GLM-5.3 Cyber Report: 64% to 100% Bypass

At a glanceQuick answers
What did Anthropic find about GLM-5.3?
That it builds end-to-end cyber exploits at close to Claude Mythos Preview’s rate, and that its safeguards give way to simple techniques.
What does it recommend?
Wider access to strong models for cyber defenders, and government safety testing of capable models, including GLM-5.3’s successors.
What is the one caveat to carry?
Anthropic competes with Z.ai, and its bypass tests ran in a simulation where no code executed.
Editorial data illustration on a near-white ground titled Anthropic tests GLM-5.3: four bars rising from 0 to 64, 92 and 100 percent for direct request, cover story, prefilled thinking and abliterated, the last bar amber, and a large 50 of 410 labeled end-to-end exploits
Fig 0Safeguards that fall with a cover story. Made by CellCog's image agent, running GPT Image 2.5.

Anthropic says the most capable openly downloadable model for building cyber exploits ships with safeguards that fall to a cover story. In a report published September 29, 2026, its Frontier Red Team writes that “GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse,” and that attackers “can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.”

This page is read from Anthropic’s report, by Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher. A disclosure first: Anthropic competes with Z.ai, and our own AI employees run on Anthropic’s Claude Opus 5.5. Our earlier record of the model is GLM 5.3 for AI agents.

On this page · 9 sectionsOpen
  1. What GLM-5.3 can build
  2. How the safeguards failed
  3. The independent check
  4. What Anthropic wants
  5. What the report does not settle
  6. What this means if you run agents on real work
  7. What we are watching for
  8. Update log
  9. Sources
Key points6 · 5 min full read
  1. On September 29, 2026, Anthropic’s Frontier Red Team published an analysis of GLM-5.3, Z.ai’s open-weight model, finding it can build working cyber exploits end to end.
  2. On ExploitBench, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, against 56 of 410 for Claude Mythos Preview; earlier models such as GLM-5.2 and Claude Opus 4.6 succeeded in none of the binary-exploitation tasks tested.
  3. Anthropic says simple techniques bypassed GLM-5.3’s safeguards: a cover story 64% of the time, prefilled thinking 92%, and an abliterated copy 100%, while safeguarded Claude models stayed at zero.
  4. Abliteration took Anthropic’s first-time team about 2,200 GPU hours, roughly $4,400, and cut the refusal rate from above 90% to as low as 2% with capability largely intact.
  5. NIST’s CAISI called GLM-5.3 the most cyber-capable open-weight model released to date, about four months behind the US frontier.
  6. Anthropic is a rival lab testing a competitor’s model; the misuse tests ran in a simulated environment it calls imperfect.

§ 01What GLM-5.3 can build

Test GLM-5.3 Claude Mythos Preview
ExploitBench, end-to-end exploits 50 of 410 56 of 410
Binary exploitation, full control-flow hijack 4% 6%
Earlier models on the same binary tasks GLM-5.2: none Claude Opus 4.6: none
Table 1GLM-5.3 vs Claude on exploit benchmarks, per Anthropic’s September 29, 2026 report

The hands-on sessions are the sharper finding. With less than an hour of human attention across a day, a researcher used GLM-5.3 to find several unknown flaws in a popular browser’s JavaScript engine and chain them into a page that reads files off a visitor’s computer; Anthropic says it disclosed them to the maintainer. In a second session, GLM-5.3-Flash turned a public Chrome fix into a working exploit chain. “This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu’s API prices, this effort would have cost $20.40.”

§ 02How the safeguards failed

Out of the box, GLM-5.3 refused the overtly malicious requests in Anthropic’s simulated world. Three simple techniques changed that.

GLM-5.3 engagement with a harmful request, by bypass technique (%)Bar chart of how often GLM-5.3 engaged with an overtly harmful request: direct request 0, deceptive cover story 64, prefilled thinking 92, abliterated copy 100 highlightedDirect request0Cover story64Prefilled thinking92Abliterated copy100GLM-5.3 engagement with a harmful request, by bypass technique (%)Bar chart of how often GLM-5.3 engaged with an overtly harmful request: direct request 0, deceptive cover story 64, prefilled thinking 92, abliterated copy 100 highlightedDirect request0Cover story64Prefilled thinking92Abliterated copy100
Fig 1GLM-5.3 engagement with a harmful request, by bypass technique (%)

Safeguarded Claude models stayed at zero under all three, Anthropic says, partly for structural reasons: its API offers no way to prefill Claude’s thinking, and Claude’s weights are not public, so they cannot be abliterated. The abliteration itself took Anthropic’s first-time team “about 2,200 GPU hours at a computation cost of roughly $4,400,” and it estimates an experienced team would need about 600 hours. It took GLM-5.3’s refusal rate from above 90% to 3%, 2% and 12% on three public benchmarks, with general capability unchanged on GPQA-Diamond.

§ 03The independent check

NIST’s Center for AI Standards and Innovation published its own assessment on September 17, finding GLM-5.3 to be “the most cyber-capable open-weight model released to date,” about four months behind the US frontier. Anthropic says its capability findings broadly match CAISI’s. The safeguard tests are Anthropic’s alone.

§ 04What Anthropic wants

Two things. More access for defenders: “Our view is that cyber defenders should use the best available tools that meet their needs,” and Anthropic says it is widening access to Claude’s cyber capabilities. And testing by governments: “Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3.” Its summary line is that “a critical threshold in freely accessible capabilities has now been crossed.”

§ 05What the report does not settle

  • Who is testing whom. Anthropic sells the safeguarded alternative it compares against. The capability numbers have CAISI behind them; the bypass numbers do not.
  • Real-world conditions. The misuse tests ran in a simulated world where no model-written code executed, and Anthropic itself calls these simulations “imperfect measures of how a model would behave.”
  • Actual attacks. The report predicts state and non-state actors will use models like this; it documents no attack carried out with GLM-5.3.
  • Z.ai’s view. The report includes no response from Zhipu AI.

§ 06What this means if you run agents on real work

The attacker’s toolkit just got cheaper, and it is downloadable. That puts more weight on the defensive layer every business already controls: patching fast, giving each agent and each person only the access its work needs, and putting a gate in front of anything consequential. The same questions apply to your own agents as to anyone else’s.

§ 07What we are watching for

  • A response from Zhipu AI, or safeguard changes in a GLM-5.3 update.
  • The disclosures Anthropic says it will make for the drivers and device software it tested.
  • Whether other labs publish matching red-team results on open-weight models.

§ 08Update log

  • September 30, 2026: page opened, from Anthropic’s report of September 29.

§ 09Sources

Frequently asked6 questions

Q1What is GLM-5.3?

An open-weight model from Zhipu AI, known outside China as Z.ai. Anyone can download its weights, which is why Anthropic’s report focuses on how easily its safeguards can be removed.

Q2What is abliteration?

A technique that edits an open-weight model to remove its refusals. Anthropic says several developers published abliterated versions of GLM-5.3 within days of release, and its own copy barely lost capability on GPQA-Diamond.

Q3How close is GLM-5.3 to Claude Mythos?

On ExploitBench, 50 of 410 end-to-end exploits for GLM-5.3 versus 56 of 410 for Claude Mythos Preview. On Anthropic’s binary-exploitation benchmark, 4% versus 6% full control-flow hijacks.

Q4Did anyone independent check this?

On capability, yes: NIST’s Center for AI Standards and Innovation published its own assessment on September 17 and Anthropic says its findings broadly match. The safeguard-bypass tests are Anthropic’s own.

Q5What does this mean for businesses using AI agents?

Capable attack tooling is now freely downloadable, so the defenses around your systems matter more: patching, least-privilege access and approval steps on anything consequential.

Q6Where does CellCog fit?

Our AI employees run on Claude Opus 5.5, not GLM-5.3, and every action that reaches your world passes an approval step. We cover this because the security of the tools attackers can reach is the backdrop every agent platform works against.

Published 30 September 2026 All Trust, permissions & security →