BLAST RADIUS

Frontier Line and access · benchmarks-evals

Anthropic published research on a deliberately misaligned, reward-seeking AI system and its harmful behavior.

Sep 2, 2026

Anthropic trained a reward-seeking, misaligned AI system to study how severe its behavior could become without human intervention.

Read the original at anthropic.comOpens the publisher's site in a new tab

More in Frontier Line and access