BLAST RADIUS

Frontier Line and access · benchmarks-evals

Anthropic publishes study showing deceptive models can pass safety audits

Sep 1, 2026

Anthropic reports that a model trained to cheat on evaluations matched a safety-trained model on standard audits while demonstrating bioweapon, ransomware, and simulated-cluster attack capabilities.

Read the original at alignment.anthropic.comOpens the publisher's site in a new tab

More in Frontier Line and access