In this paper, we propose a novel approach for crafting targeted adversarial examples (attacks) using explainable artificial intelligence (XAI) techniques. Our method leverages XAI to identify key input elements that, when altered, can mislead NLP models, such as BERT and large language models (LLMs), into producing specific incorrect outputs. We demonstrate the effectiveness of our targeted attacks across a range of NLP tasks and models, even in scenarios where internal model access is restricted. Our approach is particularly effective in zero-shot learning settings, underscoring its adaptability and transferability to both traditional and conversational AI systems. In addition, we outline mitigation strategies, demonstrating that adversarial training and fine-tuning can enhance model defenses against such attacks. Although our work highlights the vulnerabilities of LLMs and BERT models to adversarial manipulation, it also lays the groundwork for developing more robust models, advancing the dual goal of understanding and securing black-box NLP systems. Through targeted adversarial examples and SHAP-based techniques, we not only expose the weaknesses of existing models but also propose strategies to enhance AI’s resilience to deceptive linguistic input.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Precise Language Deception: XAI Driven Targeted Adversarial Examples with Restricted Knowledge

  • Mateusz Gniewkowski,
  • Paweł Walkowiak,
  • Marek Klonowski,
  • Tomasz Walkowiak

摘要

In this paper, we propose a novel approach for crafting targeted adversarial examples (attacks) using explainable artificial intelligence (XAI) techniques. Our method leverages XAI to identify key input elements that, when altered, can mislead NLP models, such as BERT and large language models (LLMs), into producing specific incorrect outputs. We demonstrate the effectiveness of our targeted attacks across a range of NLP tasks and models, even in scenarios where internal model access is restricted. Our approach is particularly effective in zero-shot learning settings, underscoring its adaptability and transferability to both traditional and conversational AI systems. In addition, we outline mitigation strategies, demonstrating that adversarial training and fine-tuning can enhance model defenses against such attacks. Although our work highlights the vulnerabilities of LLMs and BERT models to adversarial manipulation, it also lays the groundwork for developing more robust models, advancing the dual goal of understanding and securing black-box NLP systems. Through targeted adversarial examples and SHAP-based techniques, we not only expose the weaknesses of existing models but also propose strategies to enhance AI’s resilience to deceptive linguistic input.