HASTIKA: hate speech and target identification in Kannada-English code-mixed text
摘要
In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional languages, leading to code-mixed text. While this linguistic flexibility fosters freedom of expression, it also contributes to the rampant spread of hate speech, posing significant societal challenges. Identifying hate speech in low-resource Kannada-English code-mixed text is challenging due to the scarcity of annotated corpora. To address this gap, the authors introduce HASTIKA (Hate Speech and Target Identification in Kannada-English Code-Mixed Text), a gold-standard corpus specifically designed for hate speech detection and target identification. It consists of 8,058 YouTube comments, annotated for binary classification “Hate" and “Non-hate" and fine-grained categorization into “Gender", “Political", “Religion", “Geo-political", “Violence", and “Others", marking a significant contribution as the first dataset exclusively tailored for this purpose. Topic modeling techniques were applied to uncover latent themes. Benchmark experiments show that fastText, which captures word-level features, achieves 0.7529 accuracy for binary classification and 0.6042 for multi-class classification, while BERT, excelling in sentence-level feature extraction, attains 0.8054 and 0.6819 accuracy, respectively. This research provides a foundational resource for hate speech detection in Kannada-English code-mixed text and underscores the importance of linguistic and contextual information in developing robust classification models.