Robustness of AraBERT models to black-box adversarial deletions and perturbations
摘要
Large Language Models (LLMs) are becoming vulnerable to adversarial attacks due to the creation crafted inputs that may deceive even the most advanced protections. These models can produce the wrong prediction with slight changes in the input data. This flaw can be noticed in many areas, such as computer vision, speech recognition, and natural language processing, and there are serious doubts about the robustness of AraBERT LLMs in sensitive use. Consequently, this paper assesses the robustness of AraBERT LLMs to character-level and word-level black-box attacks that cause spelling errors in the input data. The adversarial samples were generated using the Chain-of-Thought prompting method, to evaluate the robustness of a fine-tuned AraBERT model and the original AarBERT v2.0 LLM. A considerable reduction in accuracy was found, especially with regard to word-level delete attacks, and the largest reduction of 44.93% of the original model was obtained. The deletions at the word and character levels were also significant. This highlights how important it is to strengthen the model’s robustness to deletions, in order to obtain reliable performance for a variety of text lengths and challenging crafted input.