<p>In modern multimedia systems, adversarial text attack is a vital way to expose the vulnerability of deep neural networks and improve their robustness. However, existing methods have some limitations. For example, character-level insertion attacks cause misspelling errors and word-level attacks tend to make limited lexical variations. Although sentence-level attacks can greatly enrich the variety of sentences, they are less effective towards fooling victim models and sometimes lead to the wrong representation. In this paper, we propose the Parentheses Insertion Sentence-level Text Adversarial Attack (PI) algorithm that crafts adversarial texts by filling frequently used parentheses. Specifically, we collect a parentheses set (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1678_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="30" /> </InlineMediaObject> <EquationSource Format="TEX">\(P_{set}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>P</mi> <mrow> <mi mathvariant="italic">set</mi> </mrow> </msub> </math></EquationSource> </InlineEquation>) at the beginning where all the parentheses are meaningless to ensure the semantics of the sentence remain unchanged after the insertion. Then we utilize the beam search strategy to merge the selected parentheses in the appropriate text positions to improve the attack success rate (ASR). To evaluate the effectiveness of PI method, we conduct extensive experiments by attacking several popular models. Experimental results show that PI enhances the ASR performance compared to word-level and sentence-level baselines while preserving high semantic similarity and incurring minimal perturbation costs. Additionally, PI helps enhance the robustness of modern NLP models by adversarial training.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parentheses insertion based sentence-level text adversarial attack

  • Ang Li,
  • Xinghao Yang,
  • Baodi Liu,
  • Honglong Chen,
  • Dapeng Tao,
  • Weifeng Liu

摘要

In modern multimedia systems, adversarial text attack is a vital way to expose the vulnerability of deep neural networks and improve their robustness. However, existing methods have some limitations. For example, character-level insertion attacks cause misspelling errors and word-level attacks tend to make limited lexical variations. Although sentence-level attacks can greatly enrich the variety of sentences, they are less effective towards fooling victim models and sometimes lead to the wrong representation. In this paper, we propose the Parentheses Insertion Sentence-level Text Adversarial Attack (PI) algorithm that crafts adversarial texts by filling frequently used parentheses. Specifically, we collect a parentheses set ( \(P_{set}\) P set ) at the beginning where all the parentheses are meaningless to ensure the semantics of the sentence remain unchanged after the insertion. Then we utilize the beam search strategy to merge the selected parentheses in the appropriate text positions to improve the attack success rate (ASR). To evaluate the effectiveness of PI method, we conduct extensive experiments by attacking several popular models. Experimental results show that PI enhances the ASR performance compared to word-level and sentence-level baselines while preserving high semantic similarity and incurring minimal perturbation costs. Additionally, PI helps enhance the robustness of modern NLP models by adversarial training.