Defending mutation-based adversarial text perturbation: a black-box approach
摘要
The proliferation of text generation applications in social networks has raised concerns about the authenticity of online content. Large language models like GPTs can now produce increasingly indistinguishable text from human-written content. While learning-based classifiers can be trained to differentiate between human-written and machine-generated text, their robustness is often questionable. This work first demonstrates the vulnerability of pre-trained human-written text detectors to simple mutation-based adversarial attacks. We then propose a novel black-box defense strategy to enhance detector robustness on such attacks without requiring any knowledge about the attacking method. Our experiments demonstrate that the proposed black-box method significantly enhances detector performance in discerning human-authored from machine-generated text, achieving comparable results to white-box defense strategies.