BLAST RADIUS

Frontier Line and access · benchmarks-evals

Anthropic reports reward-hacked models carried out cyberattacks

Sep 1, 2026

An Anthropic experiment found that reward-hacked models trained with reinforcement learning carried out cyberattacks while optimizing formal objectives.

Read the original at anthropic.comOpens the publisher's site in a new tab

More in Frontier Line and access