Social media and digital platforms play a crucial role in user’s daily lives which connects people all over the world and provides an opportunity to openly express opinions. However, as bullying has increased and abusive comments on social media can affect victims with severe depression and suicidal tendencies. Automatic abusive language detection is the key to countering these undesirable issues. The English language has been the subject of extensive research, yet Bengali-English code-mixing data or Romanized Bengali for abusive sentence classification are less explored. Here, in this study, a large, annotated Romanized Bengali-English code-mixed corpus was prepared, rigorous preprocessing was done, and also American Soundex-based word correction technique was implemented. Various Machine Learning (LR, SVM, MNB) and Deep Learning (RNN, GRU, LSTM, Bi-LSTM) models were employed to categorize comments according to their level of abuse. All these models were evaluated, and the results were compared. Bi-LSTM achieved an accuracy of 90.06%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Abusive Language Detection on Romanized Bengali and Bengali-English Code-Mixed Dataset by Deep Learning Models

  • Joyjit Mandal,
  • Ravi Raj Choudhary,
  • Gaurav Meena,
  • Manisha Barman

摘要

Social media and digital platforms play a crucial role in user’s daily lives which connects people all over the world and provides an opportunity to openly express opinions. However, as bullying has increased and abusive comments on social media can affect victims with severe depression and suicidal tendencies. Automatic abusive language detection is the key to countering these undesirable issues. The English language has been the subject of extensive research, yet Bengali-English code-mixing data or Romanized Bengali for abusive sentence classification are less explored. Here, in this study, a large, annotated Romanized Bengali-English code-mixed corpus was prepared, rigorous preprocessing was done, and also American Soundex-based word correction technique was implemented. Various Machine Learning (LR, SVM, MNB) and Deep Learning (RNN, GRU, LSTM, Bi-LSTM) models were employed to categorize comments according to their level of abuse. All these models were evaluated, and the results were compared. Bi-LSTM achieved an accuracy of 90.06%.