Classification of Political Disinformation Texts in Spanish: A Multiclass Approach Based on Machine Learning and Linguistic Features
摘要
Disinformation and online fake news represent a sociotechnical problem with significant impacts on political contexts and discourse, particularly due to the creation of multimodal false information and its social consequences. This phenomenon, studied across various fields, faces challenges such as data scarcity, limited research in multilevel classification, and insufficient studies in Spanish. This study focuses on identifying the lexical-grammatical features of disinformation texts in Spanish, contributing to the discussion within the Chilean political context by combining lexical-grammatical analysis with machine learning tools. A dataset of disinformation texts related to Chilean politics was constructed, and lexical-grammatical features were automatically extracted using six linguistic variables for their characterization. Subsequently, machine learning models were applied in both binary and multiclass classification tasks. The results revealed 22 statistically significant features (p-value \(< 0.05\) ) using ANOVA. In the multiclass classification, the best F1-score was observed in the “true” category with the KNN model (0.40). In the binary classification, Linear SVM and Naive Bayes stood out, both achieving a precision of 0.75 and an F1-score of 0.62 in the “false” category, values considered superior to those of human precision.