Cyberbullying has emerged as a significant concern in the modern world. In Bangladesh receiving hate comments and bullying on social media platforms, particularly on Facebook, has unfortunately become a common occurrence. As a low-resource language, comparatively few studies have been conducted on Bangla cyberbullying detection. This paper explores the performance of pre-trained transformer models Bangla BERT Base and Mixed Distil BERT for classifying Bangla cyberbullying texts. Our research focused on utilizing a customized-balanced dataset containing 10,000 Bangla Facebook comments collected from a publicly available dataset labelled with five classes. Our findings reveal an accuracy of 78.25% with Bangla BERT Base and 76.2% with Mixed Distil BERT. The precision, recall and F1 score for Bangla BERT Base were 78.74, 78.25 and 78.39% respectively while for Mixed Distil BERT, they were 76.43, 76.2 and 76.25% respectively. The close values of precision, recall and F1 score imply a balanced performance across different classes and provide confidence that it can handle real-world data without significant biases. These results outperform all the previous transformer-based approaches for determining five cyberbullying classes on a balanced dataset in Bengali.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cyberbullying Detection in Bangla Facebook Comments Using Pre-trained Transformer Models

  • Rama Kundu Prova,
  • Sarnali Basak

摘要

Cyberbullying has emerged as a significant concern in the modern world. In Bangladesh receiving hate comments and bullying on social media platforms, particularly on Facebook, has unfortunately become a common occurrence. As a low-resource language, comparatively few studies have been conducted on Bangla cyberbullying detection. This paper explores the performance of pre-trained transformer models Bangla BERT Base and Mixed Distil BERT for classifying Bangla cyberbullying texts. Our research focused on utilizing a customized-balanced dataset containing 10,000 Bangla Facebook comments collected from a publicly available dataset labelled with five classes. Our findings reveal an accuracy of 78.25% with Bangla BERT Base and 76.2% with Mixed Distil BERT. The precision, recall and F1 score for Bangla BERT Base were 78.74, 78.25 and 78.39% respectively while for Mixed Distil BERT, they were 76.43, 76.2 and 76.25% respectively. The close values of precision, recall and F1 score imply a balanced performance across different classes and provide confidence that it can handle real-world data without significant biases. These results outperform all the previous transformer-based approaches for determining five cyberbullying classes on a balanced dataset in Bengali.