错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Telugu-English Abusive Comment Detection Using XLMRoBERTa and mBERT

  • Pingala Revanth Reddy,
  • K. V. Munawwar,
  • K. Nandhini

摘要

The proliferation of social media platforms has enabled users to express their thoughts and opinions freely, but it has also given rise to the rampant spread of abusive and offensive content. The detection and moderation of such abusive comments have become crucial for maintaining a healthy online environment. Detecting abusive comments in multilingual settings is a challenging task due to the presence of diverse languages, writing scripts, and code-mixing. This paper presents a comprehensive approach for abusive comment detection in the Telugu-English language pair, fine-tuning the state-of-the-art models for native Telugu script, Telugu sentences written in English script, and code-mixing or a combination of Telugu and English script. We leverage the power of two state-of-the-art pre-trained language models, XLMRoBERTa and mBERT, to effectively tackle this task. Our results demonstrate the efficacy of XLMRoBERTa and mBERT models in addressing these challenges in the detection of abusive language in the multilingual context in terms of accuracy, precision, recall, and F1 score. The fine-tuned mBERT gave an Accuracy of 66.21% and fine-tuned XLMRoBERTa gave an Accuracy of 70.51%.