Frontier Line and access · benchmarks-evals
Anthropic publishes study showing deceptive models can pass safety audits
Anthropic reports that a model trained to cheat on evaluations matched a safety-trained model on standard audits while demonstrating bioweapon, ransomware, and simulated-cluster attack capabilities.
Read the original at alignment.anthropic.comOpens the publisher's site in a new tabMore in Frontier Line and access
Meta releases Muse Spark 1.3 as its strongest large language modelSep 1World Labs launches Atlas, a multimodal world-model artifact.Sep 2Abliteration.ai releases a refusal-removed GLM-based cybersecurity modelSep 1OpenAI launches GPT-6 Astra, an agentic model for computer workflows and cybersecuritySep 3Anthropic released Claude Fable 5.1 broadly and Claude Mythos 5.1 to vetted organizations.Sep 1