Arabic Hate Speech Detection on Social Media Using Machine Learning
摘要
With the spread of online hate speech, in the recent decade, that threatens the safety of human beings and assaults their protected characteristics, there is an important interest in hate speech detection on social media as a real-world problem. Our work aims to detect automatically, the hateful comments using Logistic regression (LR) and Linear Support Vector Classification (Linear SVC) as machine learning algorithms with Term Frequency – Inverse Document Frequency (TF-IDF) and Bi-directional Long Short Term Memory (BI-LSTM) as a deep learning model with word embedding, implementing them on the Arabic and Tunisian Dataset named Tunisian Hate Speech and Abusive Dataset (T-HSAB), trying to exhibit the impact of NLP techniques on Arabic text classification. Linear SVC outperforms the other models with an accuracy equal to 99.75%