Anthropic: open-weight GLM-5.3 builds cyber exploits like Claude Mythos Preview — and its safeguards fall apart
Anthropic's Frontier Red Team found Zhipu's open-weight GLM-5.3 can autonomously build end-to-end cyber exploits nearly as well as Claude Mythos Preview, and its safeguards bypass at 64-100% rates.
What happened
On September 29, 2026, Anthropic's Frontier Red Team published an analysis of GLM-5.3, the open-weight model from Zhipu AI (branded Z.ai outside China). The headline result: GLM-5.3 can autonomously build complete end-to-end cyber exploits nearly as often as Claude Mythos Preview — Anthropic's own model that debuted this capability about five months ago (April 2026) and was only shared with vetted defenders through Project Glasswing. Those defenders have reportedly surfaced more than 10,000 vulnerabilities in critical software since.
On Anthropic's ExploitBench, which uses known flaws in Chrome's V8 JavaScript engine, GLM-5.3 succeeded in 50 of 410 attempts, versus 56 for Mythos Preview.
In a hands-on session, GLM-5.3 found several previously unknown flaws in a widely used browser's JavaScript engine (Linux build) in one day and chained them into a webpage that reads arbitrary files from a visitor's machine. The smaller GLM-5.3-Flash turned a disclosed Chrome flaw plus one other known bug into a reliable exploit chain with about 20 minutes of human time and 8 hours of model work — costing roughly $20.40 at Zhipu's API prices.
The safeguards fared worse. Out of the box GLM-5.3 refused openly hostile requests, but a deceptive red-team cover story raised engagement to 64%, prefilled reasoning pushed it to 92%, and abliteration — stripping refusals from the open weights — worked every time. Anthropic estimates an experienced team needs about $1,200 of compute; refusal rates fell from over 90% to 2–12%. Abliterated copies were public within days of the model's release. Anthropic notes the same techniques did not work against safeguarded Claude models in its testing.
NIST's Center for AI Standards and Innovation (CAISI) independently assessed GLM-5.3 on September 17 as the most cyber-capable open-weight model to date — about four months behind US frontier models.
Why it matters
Frontier-grade exploit development is no longer gated behind trusted-access programs — it is a free download. Anthropic concludes that state and non-state actors will likely use models like GLM-5.3 to cause real harm, and is asking governments to safety-test capable models while urging defenders to run models at least as capable as what attackers now have. The report also lands one day after the White House AI accord and the NYC Council AI-safety hearing, underscoring the policy collision.
Sources
- Report: https://www.metatalks.ai/open-weight-glm-5-3-nears-mythos-preview-at-building-exploits/
- Report: https://the-decoder.com/anthropic-says-zhipus-open-weight-glm-5-3-nearly-matches-claude-mythos-preview-at-building-exploits/
- Report: https://aidailypost.com/news/anthropic-warns-zhipus-glm-53-model
- Report: https://letsdatascience.com/news/anthropic-red-team-reports-glm-53-exploitation-result-38d8105f
Sources
- ReportMetaTalks report: open-weight GLM-5.3 nears Mythos Preview at building exploits
- ReportThe Decoder report: Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits
- ReportAI Daily Post report: Anthropic warns Zhipu's GLM-5.3 model
- ReportLet's Data Science report: Anthropic red team reports GLM-5.3 exploitation result