Recognizing Hate Speech on Twitter with Feature Combo
摘要
The issue of hate speech directed toward women is widespread and has gained increased attention in recent times. Despite the impressive performance of machine learning-based models that incorporate textual, user-specific, and social network features, there is still potential for improvement given the variety of feature combinations used with the ensemble learning (EL) approach. To fill this gap, researchers in this study have generated a unique set of features in terms of stance, and similarity, and combined with machine learning (ML), and EL algorithms to recognize hate speech in Twitter data, and assess the model’s effectiveness. The proposed novel approach of feature combo with stance and similarity features showed highest accuracy with ensemble algorithm namely Extreme Gradient Boosting (XGBoost), of 93.53%, while Support Vector Machine (SVM) algorithm of ML showed lowest accuracy of 75.67%.