The proliferation of fake news on digital platforms presents significant societal challenges, necessitating the development of robust and interpretable detection systems. This study proposes a stacking-based ensemble learning model that integrates XGBoost and Logistic Regression to improve fake news classification accuracy while enhancing model transparency. Unlike traditional Natural Language Processing (NLP) approaches, which rely solely on textual analysis, this model incorporates structural and statistical metadata features, such as publication date and article length, to improve generalizability across misinformation domains. Experimental results on the Spanish Political Fake News dataset demonstrate that the stacking ensemble model outperforms individual classifiers, achieving an F1-score of 95.2% and a ROC-AUC of 0.974. SHAP (Shapley Additive Explanations) analysis enhances interpretability by identifying the most influential features contributing to classification decisions, confirming that metadata plays a critical role in misinformation detection. These findings highlight the effectiveness of hybrid machine-learning approaches that combine textual and structural information for scalable misinformation detection. The study’s contributions include a highly accurate and explainable classification model, positioning ensemble learning as a viable solution for real-world applications in automated fact-checking, journalism, and social media moderation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensemble Learning for Fake News Detection: Enhancing Classification Accuracy and Explainability with Structural and Statistical Metadata

  • Sebastián González-Celi,
  • Henry N. Roa,
  • Jorge Cruz-Silva,
  • Edison Loza-Aguirre,
  • Nelson Salgado-Reyes,
  • Javier Guaña-Moya

摘要

The proliferation of fake news on digital platforms presents significant societal challenges, necessitating the development of robust and interpretable detection systems. This study proposes a stacking-based ensemble learning model that integrates XGBoost and Logistic Regression to improve fake news classification accuracy while enhancing model transparency. Unlike traditional Natural Language Processing (NLP) approaches, which rely solely on textual analysis, this model incorporates structural and statistical metadata features, such as publication date and article length, to improve generalizability across misinformation domains. Experimental results on the Spanish Political Fake News dataset demonstrate that the stacking ensemble model outperforms individual classifiers, achieving an F1-score of 95.2% and a ROC-AUC of 0.974. SHAP (Shapley Additive Explanations) analysis enhances interpretability by identifying the most influential features contributing to classification decisions, confirming that metadata plays a critical role in misinformation detection. These findings highlight the effectiveness of hybrid machine-learning approaches that combine textual and structural information for scalable misinformation detection. The study’s contributions include a highly accurate and explainable classification model, positioning ensemble learning as a viable solution for real-world applications in automated fact-checking, journalism, and social media moderation.