An increasing number of users utilize public platforms to communicate. The languages used by general public are diverse and varied. Detection of the offensive words utilized by people on these online platforms is challenging. Unstructured texts, inappropriate models combined with low-resource datasets, cause the research on less prevalent languages quite arduous. In this chapter, we have attempted to address the problem of offensive comments detection in Bengali, which is a low-resource language with limited approaches. Our approach is a binary classification model that successfully classifies toxic comments from nontoxic comments. Natural Language Processing techniques have been implemented for preprocessing of the unstructured data. Deep learning model with a modified version of BERT architecture has been applied. This model has been pretrained on a large corpus of Bengali text and efficiently detects the offensive words in the social media Bengali comments. The proposed approach is performing effectively when compared with different baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Toxic Comments in Bengali Language

  • Abir Mondal,
  • Kingshuk Roy,
  • Susmita Das,
  • Arpita Dutta

摘要

An increasing number of users utilize public platforms to communicate. The languages used by general public are diverse and varied. Detection of the offensive words utilized by people on these online platforms is challenging. Unstructured texts, inappropriate models combined with low-resource datasets, cause the research on less prevalent languages quite arduous. In this chapter, we have attempted to address the problem of offensive comments detection in Bengali, which is a low-resource language with limited approaches. Our approach is a binary classification model that successfully classifies toxic comments from nontoxic comments. Natural Language Processing techniques have been implemented for preprocessing of the unstructured data. Deep learning model with a modified version of BERT architecture has been applied. This model has been pretrained on a large corpus of Bengali text and efficiently detects the offensive words in the social media Bengali comments. The proposed approach is performing effectively when compared with different baselines.