<p>With the rapid development of artificial intelligence technology, Large Language Models (LLMs) have become the focus of academic and industrial attention. For example, OpenAI’s ChatGPT can simulate human writing styles and automatically construct text such as news stories, academic reports, and technical documents for use in a wide range of applications, including automated content creation, customer service, and educational assistance. However, the widespread use of such texts also brings a series of problems, such as the spread of false information, academic integrity issues, and the impact on people’s ability to think independently. The progress of LLMs in text generation ability makes the text generated by LLMs difficult to distinguish from the text written by humans, which affects the authenticity of information, academic integrity and the reliability of text content. Therefore, in order to deal with these challenges, effective detection of machine-generated text is crucial. Existing machine-generated text detection methods mainly use the statistical characteristics of the text or train classifiers to detect such text. However, the existing methods have certain shortcomings in learning deep features to distinguish human text from machine text, which limits their detection performance. This paper proposes a machine-generated text detection model GenDetect-GAN based on generative adversarial networks to solve this problem. The generator learns the distribution characteristics of real human text data and generates highly realistic text samples. The discriminator is continuously trained through these samples to improve its detection performance. The experimental results show that GenDetect-GAN outperforms the existing machine-generated text detection models on multiple indicators.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GenDetect-GAN: generative adversarial network model for machine-generated text detection

  • Bin Xu,
  • Wenjun Yu,
  • Qin Wen,
  • Qiulan Cui,
  • Jin Qi

摘要

With the rapid development of artificial intelligence technology, Large Language Models (LLMs) have become the focus of academic and industrial attention. For example, OpenAI’s ChatGPT can simulate human writing styles and automatically construct text such as news stories, academic reports, and technical documents for use in a wide range of applications, including automated content creation, customer service, and educational assistance. However, the widespread use of such texts also brings a series of problems, such as the spread of false information, academic integrity issues, and the impact on people’s ability to think independently. The progress of LLMs in text generation ability makes the text generated by LLMs difficult to distinguish from the text written by humans, which affects the authenticity of information, academic integrity and the reliability of text content. Therefore, in order to deal with these challenges, effective detection of machine-generated text is crucial. Existing machine-generated text detection methods mainly use the statistical characteristics of the text or train classifiers to detect such text. However, the existing methods have certain shortcomings in learning deep features to distinguish human text from machine text, which limits their detection performance. This paper proposes a machine-generated text detection model GenDetect-GAN based on generative adversarial networks to solve this problem. The generator learns the distribution characteristics of real human text data and generates highly realistic text samples. The discriminator is continuously trained through these samples to improve its detection performance. The experimental results show that GenDetect-GAN outperforms the existing machine-generated text detection models on multiple indicators.