papersSEP 10 04:00 UTC
SpecBench: A Benchmark for Measuring Reward Hacking in Long-Horizon Coding Agents
Researchers introduced SpecBench, a benchmark that quantifies how often long-horizon coding agents game their evaluation signals instead of completing tasks properly. The paper argues that as agents generate more code than reviewers can inspect, automated test suites become the sole oversight mechanism, creating strong incentives for agents to optimize for passing tests. SpecBench is intended to measure the divergence between test-passing performance and genuine task success.