Bengali Hate Speech Detection with BERT and Deep Learning Models
摘要
An increasing amount of harmful effects have been linked to prolonged exposure to abusive language on numerous social media sites. If we want to keep the internet safe and peaceful, we must do something about the epidemic of harsh language. Although studies on the topic of identifying hostile speech have been conducted, the vast majority have only covered the English language. Recent instances in Bangladesh, however, have led to the emergence of inflammatory speech in a variety of languages. Therefore, it is crucial to address this type of harmful material. Unfortunately, Bangla hate speech detection on social media sites such as Facebook and YouTube has been hampered by a lack of available public Bangla datasets. Although some datasets are available online, they are sparse, poorly sequenced, and lack necessary data types. As a means of filling this void, we have compiled a new dataset consisting of 8600 user comments from Facebook and YouTube, which we have divided into the following five categories: sports, religion, politics, entertainment, and others. Following that, we used five distinct models to perform a massive study of abusive language in Bengali. After testing a number of different models, we found that the BERT model had the highest accuracy of 80%. The availability of this dataset greatly aids our contribution to the study of identifying hate speech in Bengali. The same models have also been run on an existing dataset of 30,000 records, where we achieved an accuracy of 97%.