Despite the significant potential of textual targeted adversarial attacks in delineating class relationships and enhancing natural language model robustness, their exploration is limited. Targeted adversarial attacks on discrete data are particularly challenging due to the lack of methods that can generate target class-relevant perturbation words. We propose Autocue, an approach that leverages target class information prompts to guide masked language models in targeted text adversarial attacks. Compared to targeted attacks implemented by existing methods, Autocue can generate smooth and natural target class adversarial samples at a higher success rate. Additionally, we assess the robustness of commonly used pre-trained models through targeted attacks and analyze their inter-class adversarial robustness. The source code is available at https://github.com/cgly/AutoCue .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Autocue : Targeted Textual Adversarial Attacks with Adversarial Prompts

  • He Zhu,
  • Ce Li,
  • Haitian Yang,
  • Yan Wang,
  • Weiqing Huang

摘要

Despite the significant potential of textual targeted adversarial attacks in delineating class relationships and enhancing natural language model robustness, their exploration is limited. Targeted adversarial attacks on discrete data are particularly challenging due to the lack of methods that can generate target class-relevant perturbation words. We propose Autocue, an approach that leverages target class information prompts to guide masked language models in targeted text adversarial attacks. Compared to targeted attacks implemented by existing methods, Autocue can generate smooth and natural target class adversarial samples at a higher success rate. Additionally, we assess the robustness of commonly used pre-trained models through targeted attacks and analyze their inter-class adversarial robustness. The source code is available at https://github.com/cgly/AutoCue .