Evaluating the adversarial robustness of Arabic spam classifiers
摘要
Several studies have exposed the vulnerability of Natural Language Processing (NLP) models to adversarial attacks, which are inputs crafted by attackers to deceive NLP models. Adversarial robustness measures the performance of these systems under such attacks. In the Arabic literature, very limited contributions exist in the field of NLP adversarial robustness. Hence, this study focuses on examining the adversarial robustness of an Arabic NLP model, especially in a classic black-box spam evasion scenario. This work introduces eight diverse adversarial attacks (character, word, sentence, and multi-level) against Arabic NLP models. Moreover, we employ local post-hoc explanations such as SHapely Additive exPlanations (SHAP) to optimize the attacks strategies. Despite the excellent unfortified model’s accuracy of 99.4%, three of the attacks reduced the accuracy by more than 90%. This work also employs local post-hoc explanations such as SHapely Additive exPlanations (SHAP) Nevertheless, the proposed adversarial attacks were effective in terms of maintaining high semantic similarity and low perturbation distance.