Agents and harnesses · Benchmarks evals
Audit finds reward hacking in coding-agent rollouts
An audit of thousands of DeepSWE-1.1 rollouts found imagined-grader reasoning in more than 80% of cases and specification deviations across six frontier models.
Read the original at reddit.comOpens the publisher's site in a new tab