错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Classification of Human and Machine-Generated Texts Using Lexical Features and Supervised/Unsupervised Machine Learning Algorithms

  • Jonathan Rojas-Simón,
  • Yulia Ledeneva,
  • René Arnulfo García-Hernández

摘要

In today’s digital information era, distinguishing between human- and machine-generated texts has become a focus of study in academia and industry. This is because Large-Language Models (LLMs) can produce high-quality texts, posing a challenge to the legitimacy and authenticity of texts. In this regard, it is essential to create methods and models that can differentiate whether a human or an LLM wrote a text. Therefore, this paper explores the effectiveness of supervised and unsupervised machine learning algorithms using lexical features. Mainly, we focused on traditional algorithms, such as Multilayer Perceptron (MLP), Naive Bayes (NB), Logistic Regression (LR), Agglomerative Hierarchical Clustering (AHC), and K-means Clustering (KC). Obtained results have been compared to state-of-the-art approaches presented in the Automated Text Identification (AuTexTification) shared task, serving as reference methods. Moreover, we have found that both NB and KC may achieve competitive results in the before-mentioned task.