Dotsfeed
← News

Anthropic gives Claude Sonnet 5.5 flagship cyber safeguards — and routes risky requests to the older Sonnet 5

Verified· Sep 30, 2026Published Sep 30, 2026

Claude Sonnet 5.5 is the first Sonnet to ship with Opus-class cyber safeguards; higher-risk requests visibly fall back to Sonnet 5, and vetted defenders can apply for tiered access.

What happened

Anthropic's Sept 28 launch of Claude Sonnet 5.5 ships with cybersecurity safeguards the company previously reserved for its most capable models — the first Sonnet to get them. Anthropic says Sonnet 5.5's cyber skills are a "large improvement" over Sonnet 5 and comparable to Opus 5. Higher-risk requests (penetration testing, exploit generation) are detected and visibly handed to the older Sonnet 5; routine bug fixing is unaffected, though Anthropic warns of increased refusals on benign cybersecurity tasks.

System-card numbers: with safeguards off, Sonnet 5.5 averaged 11.53 capability flags and an 80% capture rate on ExploitBench, produced 178 full arbitrary-code-execution exploits (vs 1 for Sonnet 5), completed 46.1% of CyScenarioBench challenges (vs 0.7%), and 50 control-flow hijacks on the Binary Exploitation Benchmark (vs 3). Safeguards run in three stages (probe on internal activations, lightweight on-model classifier, separate trained LLM classifier) with 99.43% recall and a 21.0% attack success rate on the rewind-attacker robustness evaluation, improved from 57.2% for Sonnet 5. API developers must opt in to automatic fallback.

Why it matters

The industry's dangerous cyber capability is moving down-market in weeks — this is the first time flagship-level guards ship on a mid-tier model. Anthropic is shifting from limiting what models can do to gating who can use the sharpest cyber skills: an expanded Cyber Verification Program gives vetted defenders tiered access to advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos. Sonnet 5.5 is also the first Sonnet with anti-distillation classifiers and expanded "preserved thinking" so reasoning cannot be decoupled from the account. Watch Haiku 5.5 — if the cheapest tier also needs cyber gates, "frontier" will no longer mark where AI risk sits.

Sources

  1. TheStreet: Anthropic Claude Sonnet 5.5 Cybersecurity
  2. VentureBeat: Anthropic Launches Claude Sonnet 5.5
  3. Unite.AI: Anthropic Releases Claude Sonnet 5.5
  4. FoneArena: Anthropic Claude Sonnet 5.5 Features

Verified September 30, 2026.

Sources

Anthropic gives Claude Sonnet 5.5 flagship cyber safeguards — and routes risky requests to the older Sonnet 5 · Dotsfeed