Detecting bots on social media platforms is a major challenge, as these automated entities are constantly evolving to evade detection. In this study, we investigate the main features that contribute to the difficulty of bot detection. Leveraging the TwiBot-20 dataset, we analyze the characteristics of misclassified accounts and explore the reasons behind their erroneous classification. Our approach combines feature engineering, Machine Learning with Random Forest, and the interpretation of model predictions using SHAP (SHapley Additive exPlanations) values. We employ clustering techniques to identify patterns in feature contributions and provide insights into the complexities of distinguishing between human and automated accounts. Our findings highlight the nature of bot detection and the need for advanced methods to address the problem of social media manipulation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Why a Bot is Undetectable? An Explainability-Based Study of Misclassified Automated Accounts in Social Networks

  • Salvador López-Joya,
  • José Ángel Díaz-García,
  • María Dolores Ruiz,
  • María José Martín-Bautista

摘要

Detecting bots on social media platforms is a major challenge, as these automated entities are constantly evolving to evade detection. In this study, we investigate the main features that contribute to the difficulty of bot detection. Leveraging the TwiBot-20 dataset, we analyze the characteristics of misclassified accounts and explore the reasons behind their erroneous classification. Our approach combines feature engineering, Machine Learning with Random Forest, and the interpretation of model predictions using SHAP (SHapley Additive exPlanations) values. We employ clustering techniques to identify patterns in feature contributions and provide insights into the complexities of distinguishing between human and automated accounts. Our findings highlight the nature of bot detection and the need for advanced methods to address the problem of social media manipulation.