SamPar: A Marathi Hate Speech Dataset for Homophobia, Transphobia
摘要
Marathi-speaking communities, especially those experiencing a change of heart after Article 377 was repealed in India, have expressed their sentiments regarding the LGBTQ+ community on social media. Leveraging a meticulously curated dataset extracted from prominent social media platforms, YouTube and Facebook, the study unveils the social, cultural, and moral perspectives palpable across both urban and rural domains. The data, derived via a rigorous manual scraping methodology, was categorized into ‘Homophobic’, ‘Non-LGBTQ+’, and ‘Transphobic’ comments, attaining a commendable Cohen’s Kappa score of 0.967, thus reflecting a high inter-annotator agreement. Our exploitative analysis unveiled a stark presence of homophobic (15.84%) and transphobic (10.52%) remarks within the digital discussions. Employing an array of machine learning and deep learning models, with a particular spotlight on Decision Trees and LSTM which achieved macro F1-scores of 0.45 and 0.53 respectively, the study not only elevates the understanding of the sociolinguistic intricacies prevalent in digital dialogues but also lays a substantive foundation for future research. It assists in understanding and potentially curtailing digital hate speech and discriminatory remarks against the LGBTQ+ community in the Marathi digital diaspora. Thus, the paper propels the conversation towards crafting a more inclusive, empathetic, and supportive digital environment, synchronizing technological prowess with socio-cultural cognizance.