COOL: Classification of Online Offensive Language Using Machine Learning and Deep Learning
摘要
In the dynamic realm of online communication, the surge in offensive language and hate speech has emerged as a critical concern. The proliferation of digital platforms has led to a distressing uptick in such behaviors, challenging the establishment of a secure and inclusive online environment. Identifying and categorizing such conduct is imperative from both an ethical standpoint and for the prevention of further harm. The training and validation dataset encompass a diverse array of online platforms, ensuring inclusivity across various communication styles and linguistic nuances. To capture the nuanced characteristics of abusive language, a range of feature extraction techniques were included, consisting of traditional NLP methods and state-of-the-art deep learning architectures. In tandem with these approaches, comprehensive experiments employing a variety of classifiers, such as logistic regression, SVM, stochastic gradient descent, decision trees, and ensemble models were conducted. In summary, this research contributes significantly to the ongoing battle against online toxicity and the promotion of more constructive online conversations. The RNN algorithm’s 99.47% accuracy rate in detecting hate speech has significant societal and platform-level ramifications. It represents a strong barrier against hate speech spreading online, creating a safer atmosphere.