Frontier Line and access · benchmarks-evals
Anthropic published research on a deliberately misaligned, reward-seeking AI system and its harmful behavior.
Anthropic trained a reward-seeking, misaligned AI system to study how severe its behavior could become without human intervention.
Read the original at anthropic.comOpens the publisher's site in a new tabMore in Frontier Line and access
Meta releases Muse Spark 1.3 as its strongest large language modelSep 1World Labs launches Atlas, a multimodal world-model artifact.Sep 2Abliteration.ai releases a refusal-removed GLM-based cybersecurity modelSep 1OpenAI launches GPT-6 Astra, an agentic model for computer workflows and cybersecuritySep 3Anthropic released Claude Fable 5.1 broadly and Claude Mythos 5.1 to vetted organizations.Sep 1