Agents and harnesses · Benchmarks evals
Specific Publishes Benchmark For Enterprise Coding Agents
Specific introduced Real-SWE, a benchmark evaluating AI models on private, real-world enterprise codebases.
Read the original at cognition.comOpens the publisher's site in a new tab