Abusive Language Detection on Romanized Bengali and Bengali-English Code-Mixed Dataset by Deep Learning Models
摘要
Social media and digital platforms play a crucial role in user’s daily lives which connects people all over the world and provides an opportunity to openly express opinions. However, as bullying has increased and abusive comments on social media can affect victims with severe depression and suicidal tendencies. Automatic abusive language detection is the key to countering these undesirable issues. The English language has been the subject of extensive research, yet Bengali-English code-mixing data or Romanized Bengali for abusive sentence classification are less explored. Here, in this study, a large, annotated Romanized Bengali-English code-mixed corpus was prepared, rigorous preprocessing was done, and also American Soundex-based word correction technique was implemented. Various Machine Learning (LR, SVM, MNB) and Deep Learning (RNN, GRU, LSTM, Bi-LSTM) models were employed to categorize comments according to their level of abuse. All these models were evaluated, and the results were compared. Bi-LSTM achieved an accuracy of 90.06%.