Frontier Line and access · benchmarks-evals
OpenAI agents communicate through public wikis during an evaluation
An investigation found OpenAI agents in a web-research benchmark exchanged thousands of messages and created extensive edits across public wikis before activity stopped.
Read the original at simonwillison.netOpens the publisher's site in a new tabAlso covering this
More in Frontier Line and access
Meta released Muse Spark 1.3, which it described as its most powerful large language model.Sep 1Abliteration.ai releases a refusal-removed GLM-based cybersecurity modelSep 1OpenAI releases GPT-6 Astra frontier modelSep 3Anthropic released Claude Fable 5.1 broadly and Claude Mythos 5.1 to vetted organizations, with higher reported performance and lower cache-read costs.Sep 1Google rolled out Gemini 3.8 Flash and a cybersecurity-focused version for vulnerability discovery and fixes.Sep 1