Backdoor attacks on language models are a technique that enables models to establish strong correlations between “poison samples” and the “target class” expected by attackers. In recent years, with the continuous development of game theory between backdoor attacks on language models and their corresponding defense techniques, various types of backdoor attack methods such as input trigger, prompt trigger, instruction trigger, and example trigger have achieved good performance. However, existing methods suffer from issues such as triggers being easily detected and low accuracy in predicting clean samples. To address these issues, this paper introduces a novel backdoor attack approach grounded in prompt learning. By employing the soft prompt template as both a trigger and a tool for optimization, our method identifies the most effective soft prompts for diverse sample categories, achieving stealthy and potent backdoor attacks. Our experiments indicate that this approach outperforms existing methods in terms of effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SoftPromptAttack: Advancing Backdoor Attacks in Language Models Through Prompt Learning Paradigms

  • Dixuan Chen,
  • Hongyang Yan,
  • Jiatong Lin,
  • Fan Chen,
  • Yu Cheng

摘要

Backdoor attacks on language models are a technique that enables models to establish strong correlations between “poison samples” and the “target class” expected by attackers. In recent years, with the continuous development of game theory between backdoor attacks on language models and their corresponding defense techniques, various types of backdoor attack methods such as input trigger, prompt trigger, instruction trigger, and example trigger have achieved good performance. However, existing methods suffer from issues such as triggers being easily detected and low accuracy in predicting clean samples. To address these issues, this paper introduces a novel backdoor attack approach grounded in prompt learning. By employing the soft prompt template as both a trigger and a tool for optimization, our method identifies the most effective soft prompts for diverse sample categories, achieving stealthy and potent backdoor attacks. Our experiments indicate that this approach outperforms existing methods in terms of effectiveness.