The increasing accessibility of generative AI has sparked concerns on academic integrity in academia. Many educators turn to AI text detectors to flag AI-generated submissions. However, previous studies have shown variability in the performance of such tools in different contexts. Furthermore, some demographics and cultures are not covered by previous empirical studies, despite existing evidence that detectors can be biased against certain writing styles. In this paper, we present an empirical study of five AI text detectors in classifying student-written and AI-generated essays written by Filipino senior high school and undergraduate students. Furthermore, we present an analysis of linguistic features of human-written and AI-generated essays in relation to distinguishing between human-written and AI-generated text. Our findings reveal that while AI detectors exhibit high accuracy in identifying AI-generated texts, there were a number of instances where human-written text was falsely flagged as AI-generated. We also found some trends on linguistic features that can have implications in detector performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Reliability of AI Text Detectors on Filipino Student Essays

  • Andres Lucio Chan,
  • Ralph Cedric Chua,
  • Dylan Del Rio,
  • Brandon Chase Lee,
  • Renzo Erick Ong,
  • Thomas James Tiam-Lee

摘要

The increasing accessibility of generative AI has sparked concerns on academic integrity in academia. Many educators turn to AI text detectors to flag AI-generated submissions. However, previous studies have shown variability in the performance of such tools in different contexts. Furthermore, some demographics and cultures are not covered by previous empirical studies, despite existing evidence that detectors can be biased against certain writing styles. In this paper, we present an empirical study of five AI text detectors in classifying student-written and AI-generated essays written by Filipino senior high school and undergraduate students. Furthermore, we present an analysis of linguistic features of human-written and AI-generated essays in relation to distinguishing between human-written and AI-generated text. Our findings reveal that while AI detectors exhibit high accuracy in identifying AI-generated texts, there were a number of instances where human-written text was falsely flagged as AI-generated. We also found some trends on linguistic features that can have implications in detector performance.