<p>The proliferation of text generation applications in social networks has raised concerns about the authenticity of online content. Large language models like GPTs can now produce increasingly indistinguishable text from human-written content. While learning-based classifiers can be trained to differentiate between human-written and machine-generated text, their robustness is often questionable. This work first demonstrates the vulnerability of pre-trained human-written text detectors to simple mutation-based adversarial attacks. We then propose a novel black-box defense strategy to enhance detector robustness on such attacks without requiring any knowledge about the attacking method. Our experiments demonstrate that the proposed black-box method significantly enhances detector performance in discerning human-authored from machine-generated text, achieving comparable results to white-box defense strategies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Defending mutation-based adversarial text perturbation: a black-box approach

  • Demetrio Deanda,
  • Izzat Alsmadi,
  • Jesus Guerrero,
  • Gongbo Liang

摘要

The proliferation of text generation applications in social networks has raised concerns about the authenticity of online content. Large language models like GPTs can now produce increasingly indistinguishable text from human-written content. While learning-based classifiers can be trained to differentiate between human-written and machine-generated text, their robustness is often questionable. This work first demonstrates the vulnerability of pre-trained human-written text detectors to simple mutation-based adversarial attacks. We then propose a novel black-box defense strategy to enhance detector robustness on such attacks without requiring any knowledge about the attacking method. Our experiments demonstrate that the proposed black-box method significantly enhances detector performance in discerning human-authored from machine-generated text, achieving comparable results to white-box defense strategies.