<p>In text classification tasks, deep neural networks are particularly susceptible to word-level adversarial attacks, where even small lexical perturbations can significantly change the predictions of the model. To address this challenge, this work proposes a novel defense framework, Gradient-based Synonym Substitution Attack - Adversarial Training (GSSA-AT), which integrates the Gradient-based Synonym Substitution Attack (GSSA) with adversarial training. Unlike other query-based methods, GSSA generates semantically consistent adversarial examples through a gradient-guided synonym substitution mechanism, while GSSA-AT leverages these examples alongside a loss function formulated in this paper to train the model and enhance its robustness. Experimental results demonstrate that GSSA-AT exhibits effective defense capabilities against various adversarial attacks, while also showing unique advantages in reducing the transferability of adversarial examples. Evaluations on three benchmark datasets reveal that GSSA-AT strikes an optimal balance between robustness and natural language understanding, ensuring the generalization of the model to clean data. These results establish GSSA-AT as an efficient and deployable defense solution for text classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GSSA-AT: a textual adversarial defense method based on synonym substitution

  • Min Zhu,
  • HuanZhou Li,
  • ZhangGuo Tang,
  • HanCheng Long,
  • Hao Yan,
  • Jian Zhang

摘要

In text classification tasks, deep neural networks are particularly susceptible to word-level adversarial attacks, where even small lexical perturbations can significantly change the predictions of the model. To address this challenge, this work proposes a novel defense framework, Gradient-based Synonym Substitution Attack - Adversarial Training (GSSA-AT), which integrates the Gradient-based Synonym Substitution Attack (GSSA) with adversarial training. Unlike other query-based methods, GSSA generates semantically consistent adversarial examples through a gradient-guided synonym substitution mechanism, while GSSA-AT leverages these examples alongside a loss function formulated in this paper to train the model and enhance its robustness. Experimental results demonstrate that GSSA-AT exhibits effective defense capabilities against various adversarial attacks, while also showing unique advantages in reducing the transferability of adversarial examples. Evaluations on three benchmark datasets reveal that GSSA-AT strikes an optimal balance between robustness and natural language understanding, ensuring the generalization of the model to clean data. These results establish GSSA-AT as an efficient and deployable defense solution for text classification.