错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Abusive Social Media Comments Detection for Tamil and Telugu

  • Mani Vegupatti,
  • Prasanna Kumar Kumaresan,
  • Swetha Valli,
  • Kishore Kumar Ponnusamy,
  • Ruba Priyadharshini,
  • Sajeetha Thavaresan

摘要

Multilingualism has added a new dimension to the issue of abusive language detection despite the increasing number of efforts to prevent abusive content from being shared on social media. When it comes to low-resource languages such as Tamil and Telugu, the difficulty is further increased by the lack of available resources. YouTube functions as both a video-sharing platform and a social media network. YouTube allows users to establish profiles and upload videos for their followers to view, like, and comment on. Users may find it offensive and detrimental to their mental health when other users post abusive comments on videos or in response to the comments of other users. It has been observed that the language used in these comments is frequently informal and multilingual, does not always correspond to the language’s formal syntactic and lexical structure, and involves code-switching. To deal with the above issues we propose to use the multilingual pre-trained embeddings and to be more specific, our strategy can be divided into selecting suitable pre-trained models and post which adapting the models using various fine-tuning techniques to the abusive comments detection task. We use the Bidirectional Encoder Representation from Transformers (BERT) as an Encoder to generate phrase representations so that we can accurately capture the precise contextual characteristics of posts. We conduct experiments on the YouTube comments data set using various multilingual models and several fine-tuning techniques. We compared and analyzed results across the above models along with multiple classical machine learning models and found that IndicBERT and MuRIL large-cased models perform well in Tamil-English and Telugu comments data sets respectively.