The rise of ChatGPT, a highly advanced generative artificial intelligence (AI) tool, has posed significant challenges for educators in maintaining academic integrity across educational environments. This paper explores methods and strategies essential for addressing this emerging issue. Specifically, we introduce an efficient transformer-based language model that leverages fine-tuning with Bidirectional Encoder Representations from Transformers (BERT) algorithms to detect AI-generated text. During preprocessing, we extracted features and standardized the text through several steps, including lowercase conversion, tokenization, stop-word removal, token embeddings, segment embeddings, position embeddings, and space elimination to ensure high data quality. Our investigation into distinguishing essays generated by large language models (LLMs) from those written by students yielded significant findings. The BERT-uncased model delivered promising results and achieves state-of-the-art results on a diverse set of domains, where it matches or exceeds the performance of strong Transformer models , with a testing accuracy of 98% and an overall accuracy of 98.89%. Moreover, incorporating external data from various AI models, such as GPT, further enhanced performance, underscoring the benefits of diverse training data. Our study highlights the importance of careful model selection, parameter tuning, and leveraging external data in detecting AI-generated text. These findings contribute to the ongoing discourse on academic integrity, providing educators with effective tools to verify essay authenticity and uphold educational standards.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting AI-Generated Text Using Fine-Tuned Transformers: A Study on Academic Integrity

  • Manish Prajapati,
  • Santos Kumar Baliarsingh,
  • Prabhu Prasad Dev,
  • Bashir Maina Saleh,
  • Jhalak Hota,
  • Manas Ranjan Biswal

摘要

The rise of ChatGPT, a highly advanced generative artificial intelligence (AI) tool, has posed significant challenges for educators in maintaining academic integrity across educational environments. This paper explores methods and strategies essential for addressing this emerging issue. Specifically, we introduce an efficient transformer-based language model that leverages fine-tuning with Bidirectional Encoder Representations from Transformers (BERT) algorithms to detect AI-generated text. During preprocessing, we extracted features and standardized the text through several steps, including lowercase conversion, tokenization, stop-word removal, token embeddings, segment embeddings, position embeddings, and space elimination to ensure high data quality. Our investigation into distinguishing essays generated by large language models (LLMs) from those written by students yielded significant findings. The BERT-uncased model delivered promising results and achieves state-of-the-art results on a diverse set of domains, where it matches or exceeds the performance of strong Transformer models , with a testing accuracy of 98% and an overall accuracy of 98.89%. Moreover, incorporating external data from various AI models, such as GPT, further enhanced performance, underscoring the benefits of diverse training data. Our study highlights the importance of careful model selection, parameter tuning, and leveraging external data in detecting AI-generated text. These findings contribute to the ongoing discourse on academic integrity, providing educators with effective tools to verify essay authenticity and uphold educational standards.