错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unified network ensemble and chunk-based feature selection for improved social bot detection

  • Jwala Sharma,
  • Samarjeet Borah

摘要

Social bots are automated accounts, designed to mimic human behavior and participate in amplifying misinformation and manipulating public opinion. They can be responsible for controlling the decision-making ability of a human being and pose a significant threat to integrity. To maintain the integrity of online discourse, there is a need of effective social bot detection. To efficiently handle large, high-dimensional datasets, the Targeted Relevance Aggregation for Chunk-based Extraction (TRACE) has been proposed, which processes data in manageable chunks and aggregates feature importance scores to ensure a consistent and comprehensive feature selection. A novel ensemble technique Learning and Unifying Multi-Objective Integrated Network Algorithm (LUMINA) has been proposed to address class imbalance and improve model performance through a combination of dynamic clustering-based sampling, adaptive feature weighting, and multi-objective optimization. Dynamic Clustering-Based Sampling partitions the data into clusters, resamples within each cluster to balance the class distribution, and combines the resampled data to create a more representative dataset. Features selected by TRACE are further refined using Adaptive Feature Weighting in the LUMINA framework, which adjusts feature importance to emphasize the most predictive features, enhancing overall model accuracy. Finally, a multi-objective optimization technique is employed to select the best-performing model by evaluating multiple metrics—accuracy, precision, recall, and F1-score—ensuring the final model is both balanced and robust across various performance criteria. The proposed TRACE feature selection combined with the LUMINA ensemble method has been rigorously evaluated using diverse validation techniques, including holdout, tenfold cross-validation, stratified cross-validation, and leave-one-subject-out cross-validation. The proposed model impressively achieves an accuracy of 97.62% with precision and recall of 97.65% and 97.62%, respectively.