Sentence-Level Adversarial Examples in Arabic
摘要
Several studies have demonstrated that deep neural networks are susceptible to adversarial attacks - altered inputs that lead to inaccurate results from DNN-based models. Various adversaries have been exposed in the fields of computer vision and natural language processing. Nevertheless, the majority of proposed attacks in the field of NLP have been focused on assessing the performance of DNNs that were trained and fine-tuned using English datasets. This study introduces the first sentence-level adversarial attacks developed to evaluate the performance of DNN classifiers trained on Arabic. In this paper, we introduce an efficient method for creating sentence-level adversarial samples against neural text classifiers. The efficacy of our approach depends on paraphrasing the most important sentence in the text. Our findings indicate that a single paraphrased sentence is sufficient to deceive state-of-the-art neural classifiers that have been trained on Arabic datasets.