Frontier Line and access · Benchmarks evals
Rerun corrects unstable SWE-bench comparisons
A rerun of SWE-bench Verified Django evaluations after workflow fixes changed comparisons across local models, quantizations, and reasoning settings.
Read the original at reddit.comOpens the publisher's site in a new tab