<p>Social networking sites have become an important medium for online communication, enabling users from around the world to connect, share multimodal content, and express their feelings and opinions, regardless of the topic. As the number of users on these platforms increases, so does the amount of abusive language. Though most of the available resources created for handling this issue are in English, the problem of abusive speech is not restricted to any single language. As with other Natural Language Processing tasks, detecting abusive speech in low-resource languages poses significant challenges. In this study, we conduct a set of preliminary experiments for detecting hate speech in Indonesian social media. Geographically, Indonesia consists of several regions, each having its own language. The inhabitants tend to use a mix of their own local language and Bahasa (the official national language) to engage on social media, this posing significant challenges in detecting hate speech. Our contribution is twofold: (1) we manually filter available hate speech corpora to collect code-mixed data written in Indonesian-Javanese and Indonesian-Sundanese; and (2) we conduct an extensive set of experiments to determine the most robust model for detecting hate speech in this newly created corpus. Our results show that models trained on closely related languages perform better compared to those trained on languages with a higher linguistic distance. We also investigated the possibility of translating our code-mixed corpus to resource-rich languages and found that the machine translation models were not effective in handling the unique linguistic properties of code-mixed data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ngalawan Ujaran Sengit: hate speech detection in indonesian code-mixed social media data

  • Endang Wahyu Pamungkas,
  • Patricia Chiril

摘要

Social networking sites have become an important medium for online communication, enabling users from around the world to connect, share multimodal content, and express their feelings and opinions, regardless of the topic. As the number of users on these platforms increases, so does the amount of abusive language. Though most of the available resources created for handling this issue are in English, the problem of abusive speech is not restricted to any single language. As with other Natural Language Processing tasks, detecting abusive speech in low-resource languages poses significant challenges. In this study, we conduct a set of preliminary experiments for detecting hate speech in Indonesian social media. Geographically, Indonesia consists of several regions, each having its own language. The inhabitants tend to use a mix of their own local language and Bahasa (the official national language) to engage on social media, this posing significant challenges in detecting hate speech. Our contribution is twofold: (1) we manually filter available hate speech corpora to collect code-mixed data written in Indonesian-Javanese and Indonesian-Sundanese; and (2) we conduct an extensive set of experiments to determine the most robust model for detecting hate speech in this newly created corpus. Our results show that models trained on closely related languages perform better compared to those trained on languages with a higher linguistic distance. We also investigated the possibility of translating our code-mixed corpus to resource-rich languages and found that the machine translation models were not effective in handling the unique linguistic properties of code-mixed data.