错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identifying ChatGPT-Generated Essays Against Native and Non-native Speakers

  • Anoual El kah,
  • Ayman Zahir,
  • Imad Zeroual

摘要

Native Language Identification (NLI) consists of applying automatic methods to determine the native language (i.e., L1) of the writer of a given text. In the last few years, the NLI has gained more attention from scholars in several fields, such as authorship profiling, forensic and security, and language teaching. On the other hand, using the ChatGPT as a writing aid, especially by students, raises serious concerns for the academic world about the potential misuse of AI-generated content. Therefore, in this study, we examined the performance of relevant classification models (i.e., Naïve Bayes, Support Vector Machine, and Random Forest) in identifying essays that were written by native and non-native Arabic learners as well as those were totally generated by the ChatGPT. Further, we implemented four language-independent feature extraction techniques, primarily Countvectorizer, TF-IDF, Word2vec, and Glove. The dataset used for training and evaluation includes roughly 800 short essays for each category. Our findings reveal that the SVM-based classifier with TF-IDF achieved the highest accuracy of 91.14%, and the most accurately identified essays are those written by Arabic native speakers.