Diacritical Manipulations as Adversarial Attacks in Arabic NLP Systems
摘要
Deep neural networks (DNNs) have been highly successful in natural language processing (NLP) tasks. However, they were vulnerable to adversarial attacks: subtle input modifications that lead to incorrect results. Extensive research has focused on adversarial attacks to evaluate models trained on English datasets, while Arabic remains underexplored. This paper proposes diacritical manipulation as a novel black-box token-level adversarial strategy to evaluate NLP models trained on non-diacritical Arabic datasets. By adding diacritical marks to non-diacritical Arabic input, adversarial examples are crafted to evaluate the performance of state-of-the-art DNN-based classifiers. Our results indicate that the proposed attacks significantly affect the robustness of the models, thereby emphasizing the need for resilient systems against such challenges. This study helps to bridge the gap in Arabic adversarial research, noting the urgency of dealing with language-specific features when developing robust DNN models primarily for Arabic.